Startup NeoCognition has released ApprenticeBench, the first end-to-end benchmark for computer use, continuous learning, and long-horizon agentic work on a real job. An agent is hired at a virtual construction company to process accounts payable, given a six-month invoice archive, internal reference manual, and ERP guide, then starts working with detailed manager feedback in the first month and only monthly check-ins afterward. Each of the 100 tasks was validated by professional construction accountants as realistic and solvable by a competent human with access to the same materials.

Key findings:

  • Most models including Gemini, Muse Spark, Grok, and Kimi stall at 20–25%, while Fable 5.1 scores 72% and GPT-6 Astra scores 68%. A human tester reaches 51%.
  • The gap between closed and open models is the largest the authors have seen on any benchmark: 72% (Fable 5.1) vs. 18% (Kimi K3), driven not by basic UI operations but by weaker models’ inability to learn over distance.
  • The GUI tax (performance drop when switching from API to graphical interface) hits most models at 78%, but Fable 5.1 and GPT-6 Astra actually perform better through the GUI.
  • Fable 5.1 costs $6.95 per task via API versus $18.23 through the GUI, while human cost is $7.21 per task.

Related: SkillsBench Research Shows Real Impact of Skills on LLM Agents, Stanford + MIT: Meta-Harness Optimization Boosts LLM Performance, Kimi K3 Benchmarks Challenge US Frontier Models

ApprenticeBench (NeoCognition)