TensorTest formerly Ritivel · Y Combinator W26

Frontier agents can't edit video yet. We show labs where theirs fail, and why.

TensorTest builds evaluations, RL environments and training data for AI labs, starting with video editing. Our benchmark, TimelineBench, gives an agent a real client job (raw footage, audio and a brief) and grades the finished video with tests calibrated on 2,582 blind judgments by 43 professional editors.

Best agent, of 16 tested
26.8%
of 56 editing jobs resolved (15 of 56). The average agent resolves 14.0%.
Failed runs that fail on quality alone
73%
562 of 771 pass the delivery, content and brief tests, but the edit isn't good enough.
Runs where the agent claims success
93%
including 95.5% of edits that fail a test. Agents can't grade their own cut.

Every lab has headroom

Best agent from each lab · % of 56 tasks resolved · bars show 95% CI
0%25%50%75%100% OpenAI · GPT-6 Astra: 26.8% resolved (95% CI 15.8–40.3)OpenAIGPT-6 Astra26.8% Anthropic · Claude Opus 5: 23.2% resolved (95% CI 13–36.4)AnthropicClaude Opus 523.2% xAI · Grok 4.6: 12.5% resolved (95% CI 5.2–24.1)xAIGrok 4.612.5% Google · Gemini 3.8 Flash: 10.7% resolved (95% CI 4–21.9)GoogleGemini 3.8 Flash10.7% Z.ai · GLM 5.3 Flash: 10.7% resolved (95% CI 4–21.9)Z.aiGLM 5.3 Flash10.7% DeepSeek · DeepSeek Flash: 7.1% resolved (95% CI 2–17.3)DeepSeekDeepSeek Flash7.1% Alibaba · Qwen 3.8 Max: 3.6% resolved (95% CI 0.4–12.3)AlibabaQwen 3.8 Max3.6%
The top intervals overlap, so read this as distance from 100%, not as a ranking. Source: arXiv:2609.35143, detailed results table.

Why agents fail

  • They see footage as stills and transcripts. Over half their actions go to perceiving the source, and 98% of that comes before the first render.
  • They render late. The first render lands 72–90% of the way through a run, leaving little room to revise.
  • They check for defects, not craft, so they call a flat edit finished.

Why it matters for post-training: coding agents have unit tests; editing has none, so agents grade themselves and get it wrong. Our verifier is that missing test: hard checks on delivery, content and the brief, plus a quality test that stands in for editors, checked against their judgments of all 16 agents.

What we offer labs

Available now

Private eval

We run your unreleased checkpoint or agent on TimelineBench and send a failure report: tasks resolved, near misses, and how your agent works the footage. We re-score with your own model left out of the judge panel. Results stay private.

RL environments

Containerized editing tasks with media tools, fixed inputs and the verifier exposed as a reward, packaged for your sandbox and scoped with your team.

Training & eval data

New tasks on commissioned footage with training rights, plus expert-editor trajectories, preference pairs and rubrics.

Want your next checkpoint run privately before release? Results stay private, and you get a failure report on your own models.
Email founders@ritivel.com

Team

Pavan Kalyan TankalaCEO

Trained language models at Microsoft Research. IIT Bombay B.Tech + M.Tech. Published at NeurIPS, ACL and Interspeech.

Nirmit AroraCTO

Worked on making agentic systems safe at Microsoft Research. Builds our environments and runs private evals.

Gunin GuptaCBO

IIT Bombay. Previously at Kearney. Runs lab partnerships, footage rights and our editor network.