Skip to content
Loop engineering for work that has to hold

Loops engineered
around your work.

Agency operates AI systems built around the actions, feedback, and truth of a specific domain. A reasoning model tries, measures what happened, and changes course — until the work passes, or success cannot be established.

Too variable to script. Too important to judge by the model alone.
Backed by
Day One VenturesExpeditions FundContraryStartXE2
Built by engineers from
StanfordMetaApache DataFusion
01 / The gap

Between workflows
and wishful thinking.

Why the loop matters

Fixed workflows fail when the path varies. Open-loop AI can sound right without establishing that the work holds. Consequential work needs adaptive reasoning connected to real actions, measurable feedback, and external truth.

02 / The product

Explores like an agent.
Proves like software.

What makes the loop work

A bounded domain. Purpose-built instruments. Measured state after every action. An evaluator outside the model. Stopping conditions that decide when to continue, redirect, pass, or refuse.

Test-time compute gives models room to search. Loop engineering determines what they can do, what they can observe, and when the work holds.

03 / The loop

From 3% to shipped,
one lesson at a time.

Reason → Act → Observe → Evaluate → Iterate

The loop does not trust a lucky first draft. Each action changes measurable state, external evaluation determines what failed, and that result informs the next move — 3%, 7%, 12% — until the work passes, or success cannot be established. One brief, twelve attempts, eleven changed moves. Drag the bottom edge to scrub the run.

One brief, worked to done — recorded runbrief — “calm intro · one new mechanic”
Bring us work that needs its own loop — book a discovery call →
04 / Capabilities

Where engineered loops
already work.

Each domain has different actions, feedback, and standards of truth. These show the same loop-engineering discipline applied across construction, simulation, optimization, and evaluation — engineered once, then operated continuously through Operator.

Level designer

Writes levels in your native schema, simulates each one, and reads the playthrough back. Failed pacing or reachability checks change the next layout; only levels that pass the simulation reach your editor.

90% acceptance · ~120 levels shipped

Game play

Plays your build and measures what a player would feel — difficulty, pacing, fairness. Each session’s telemetry redirects the next run, and the report lands with the traces behind it.

Coming soon

QA / chaos monkey

Drives the build toward failure, watching crash logs and state assertions. What breaks steers where it probes next, and every finding arrives as a reproducible case.

Coming soon

Query optimization

Rewrites plans and indexes, then runs them against your real benchmark. Measured latency decides what survives, and only changes that are faster on your data are handed back.

Coming soon

Model training

Proposes a configuration, trains it, and evaluates against a held-out metric. The result sets the next configuration, and the run returns the checkpoint with its evaluation history.

Coming soon

Kernel development

Rewrites kernels and compiles them against your target hardware. Real timings and correctness checks decide the next candidate; you receive the kernel that measurably won.

Coming soon
FIGURES ARE SINGLE-DEPLOYMENT RECORDS, NOT GUARANTEES · REMAINING CAPABILITIES IN PROGRESS

Bring us a domain. We’ll map the loop.

We identify a valuable class of work, locate its source of truth, and determine how a productive loop can be engineered around it — then test the first one on your real cases, at no charge.

Book a discovery call