DevOps / SRE / Platform · 27.08.2026, 15:48 UTC
“Posterity will find it ludicrous”: Sai agent hits 73% on OSWorld 2.0 performing routine (but necessary) work
| Schweregrad | info |
|---|---|
| Kategorie | DevOps / SRE / Platform |
| Quelle | The New Stack ↗ |
| Veröffentlicht | 27.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
Sai, a computer agent built by Simular, has achieved a 73% success rate on OSWorld 2.0, in a benchmark update released on Thursday. The rating is based on the 108-task benchmark, which assesses everyday, lengthy professional tasks that typically take skilled humans more than 1 hour to complete.
Simular stated that this SOTA performance “places Sai ahead” of GPT-5.6 Sol at 62.57% (as reported by OpenAI) and Opus 5 at 70.57% (as reported by Anthropic), while “Sai hit the top” at about 2/3 the cost of either.
Designed and built to focus on real-world workplace tasks and functions, rather than for unfettered throughput in the pursuit of industry accolades, Sai operates on full desktop applications and webpages, calls APIs, and writes code.
This agent combines frontier and specialist models, perceives and acts on a user’s computer via dedicated interfaces, and executes complex, real-world tasks at what the company promises is “an accessible cost for individuals” and businesses.
A computer agent built for routine (but necessary) work
Simular’s co-founder & CTO Jiachen Yang tells The New Stack that he believes computer agents built for everyone’s routine (but necessary) work – e.g. recruitment outreach, validating invoices, researching the latest news – “shouldn’t burn a hole in your pocket” just because the “underlying model was trained to solve the planet’s great unsolved open math conjectures“, or similar some pursuit designed to showcase raw engineering muscle.
“Posterity will find it ludicrous that people are still building models that way right now,” says Yang. …