01 — The setting
Start with the
real workflow.
Compare a baseline and a candidate on the same tasks, with the same tools, permissions, and budget.
Start with coding agents. Extend the checks as tools, environments, and memory evolve.
Independent AI evaluation
We’re building independent evaluation to help teams iterate toward more reliable agents and safer releases.
Explore our approachEmerging agent behaviors
We’re designing evaluations for emerging behaviors across computer-use agents, multi-agent systems, connectors, memory compaction, and recursive self-improvement (RSI) systems. Each iteration brings new questions about reliability and safety.
Our approach
Hold the conditions steady. Make the difference visible.
01 — The setting
Compare a baseline and a candidate on the same tasks, with the same tools, permissions, and budget.
Start with coding agents. Extend the checks as tools, environments, and memory evolve.
02 — The behavior
Can agents use computers reliably, respect connector permissions, coordinate across agents, and retain critical instructions through memory compaction?
Check task success and safety separately. Connect each finding to an inspectable trace.
03 — The next iteration
Learn from each iteration. Re-run failures, track regressions, and bring the evidence into the next release decision.
Keep uncertainty visible. Make the next check clear.
The evidence package
We’re designing the report around what a release owner needs to review.
Design & researchWhat improved. What regressed.
Under the same conditions.
A baseline-to-candidate comparison with the task set, environment, permissions, and budget recorded alongside the findings.
A finding is the starting point.
The trace makes it inspectable.
Failure cases linked to the actions, logs, and artifacts behind them, with reproduction steps where available.
What was tested.
What still needs an answer.
Untested behavior, incomplete runs, and limits on what the evidence can support remain part of the report.
What it took to get the result.
Across the whole run.
Our goal is more accurate cost evaluation: model, tool, and compute usage across agents, retries, and coordination. Compare cost per successful task alongside quality, with measured usage, pricing assumptions, and estimation gaps kept explicit.
Our product goals
Less setup. Shorter turnaround.
From a question to reviewable evidence.
Isolated execution. Scoped permissions.
Controlled access to sensitive data.
Recoverable runs. Traceable results.
Every attempt accounted for.
Let’s start with a real question
We’re in the design and research stage, starting with teams shipping coding agents.
Start a conversation