01
Choose representative tasks
Sample contained fixes, cross-module changes, terminal work and a task that should be declined. Avoid selecting only examples that resemble a public benchmark.
Measured signals
A concise view of reported coding, agentic and reasoning performance.
Last updated August 23, 2026
66.9
DeepSWE v1.1
28.3
Terminal-Bench 3.0
42.5
SWE-Marathon v1.1
48.2
AutomationBench
84.5%
CyberGym
Reported by Z.ai. Results may not represent performance on every real-world task.
How to read the table
The reported GLM 5.3 results cover different environments. DeepSWE and SWE-Marathon focus on software-engineering work, Terminal-Bench examines command-line agents, AutomationBench measures browser or workflow automation, and CyberGym covers authorized defensive security tasks. Their score scales are not interchangeable, so a larger number in one row does not mean that benchmark is more important.
Use the figures to decide which capabilities deserve testing in your own workflow. A repository agent also depends on its harness, retrieval strategy, tool permissions, retry policy and acceptance tests. Provider limits and model configuration can change the result even when the model name is unchanged.
For a useful evaluation, freeze a repository snapshot, write the success checks before the run and record the exact endpoint, date, tools and reasoning settings. Review the final diff, tests, unnecessary changes, latency and total credits. Compare repeated completed tasks rather than one impressive answer.
01
Sample contained fixes, cross-module changes, terminal work and a task that should be declined. Avoid selecting only examples that resemble a public benchmark.
02
Record the repository commit, model identifier, provider, context policy, agent version, tool permissions, timeouts and retry allowance.
03
Run predetermined tests and review correctness, patch scope, failed commands, recovery behavior, latency, credits and human correction time.
A score can also hide variance. Repeat each task enough times to see whether success is stable, then review the failures rather than discarding them as bad runs. For agentic work, note whether the model recognized a failed command, recovered without widening the patch and stopped when the acceptance criteria passed. Publish internal results with both the success definition and the observed limitations. That makes the evaluation useful to engineers who inherit the integration after the initial model selection. Re-run the same set after a model, provider or harness update and preserve the earlier result so a regression is visible rather than anecdotal. Assign a human owner to approve the final deployment decision and its documented limits.