Platform

A simulation engine for AI.

For Agent builders

Build high-fidelity environments around the workflows your agents are expected to handle. Compare models, test every release, uncover failure modes, and generate post-training signal without touching production.

For Enterprise

Recreate your tools, policies, data structures, and edge cases in a safe simulation. Evaluate vendors and internal agents on the work that is actually yours before they reach employees or customers.

About us

New Measure is an independent eval lab, founded by Arushi Gandhi and Abhishek E and backed by Y Combinator (W26). We offer custom evals and benchmarks as a service to application AI companies.

We believe every AI company should have benchmarks that reflect the workloads its product addresses. How do different models perform? What does a specialised harness add? Where does the agent still fail? Our benchmarks help teams answer those questions and give buyers a clearer basis for comparison.

Benchmarks built around real work

We assemble five core components: software mocks, seed data, tasks, verifiers and the agent itself, including its harness and model.

We work with each team to understand its domain, the tasks that matter and what a good outcome looks like. Our environment builders and platform handle the rest: building the environments, generating task variants, running evaluations and analysing the results.

Our work spans enterprise workflows, browser agents and tasteful content. We want to help set the standard for how agents are evaluated in their domains.