Robocurve’s RoboHarm benchmark tests GPT-6 Astra on dangerous robot commands
Robocurve released the RoboHarm benchmark to test frontier AI models on real robots. In 20 trials, GPT-6 Astra attempted 19, with 17 executions of harmful actions, while Fable 5.1 refused 20% of instructions and completed 34% overall. The results are public, highlighting safety considerations for AI in physical environments.
technology10 h ago