Frontier agents can't edit video yet. We show labs where theirs fail, and why.
TensorTest builds evaluations, RL environments and training data for AI labs, starting with video editing. Our benchmark, TimelineBench, gives an agent a real client job (raw footage, audio and a brief) and grades the finished video with tests calibrated on 2,582 blind judgments by 43 professional editors.
Every lab has headroom
Why agents fail
- They see footage as stills and transcripts. Over half their actions go to perceiving the source, and 98% of that comes before the first render.
- They render late. The first render lands 72–90% of the way through a run, leaving little room to revise.
- They check for defects, not craft, so they call a flat edit finished.
Why it matters for post-training: coding agents have unit tests; editing has none, so agents grade themselves and get it wrong. Our verifier is that missing test: hard checks on delivery, content and the brief, plus a quality test that stands in for editors, checked against their judgments of all 16 agents.
What we offer labs
Private eval
We run your unreleased checkpoint or agent on TimelineBench and send a failure report: tasks resolved, near misses, and how your agent works the footage. We re-score with your own model left out of the judge panel. Results stay private.
RL environments
Containerized editing tasks with media tools, fixed inputs and the verifier exposed as a reward, packaged for your sandbox and scoped with your team.
Training & eval data
New tasks on commissioned footage with training rights, plus expert-editor trajectories, preference pairs and rubrics.
Team
Trained language models at Microsoft Research. IIT Bombay B.Tech + M.Tech. Published at NeurIPS, ACL and Interspeech.
Worked on making agentic systems safe at Microsoft Research. Builds our environments and runs private evals.
IIT Bombay. Previously at Kearney. Runs lab partnerships, footage rights and our editor network.