UniClawBench is a bilingual, capability-driven benchmark for proactive AI agents. It evaluates agents in a closed loop with an executor, a hidden answer supervisor, and a public user simulator, ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results