NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best? (juliahub.com)
hartator 4 minutes ago [-]
It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.
giwook 45 minutes ago [-]
Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?
grim_io 53 minutes ago [-]
I'd expect google to do well here, since they were historically strong at multimodal and physics.
jespinel 53 minutes ago [-]
Nice! Is is missing Codex in the agent harnesses comparison IMO.
gizmodo59 50 minutes ago [-]
Yet another "benchmark to promote their own harness"
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 16:51:18 GMT+0000 (Coordinated Universal Time) with Vercel.