A desktop coding harness
Projects, repository intelligence, terminals, Git/worktrees, browser verification, durable threads, recovery and multiple agent runtimes in one engineering workspace.
Trebell Code is a provider-independent desktop coding harness between AI models and your repository. Run Trebell Native, Codex, Claude Code, OpenCode and more in one workspace while Trebell owns the context, tools, Git, terminals, state, recovery and verification around them. The goal is simple: reach verified results with less wasted inference.
A coding model does not ship software by itself. The harness decides what context it sees, which tools it can use, how failures recover, what gets verified, and how much inference is wasted on the way. Trebell is a neutral engineering layer around increasingly capable models.
Projects, repository intelligence, terminals, Git/worktrees, browser verification, durable threads, recovery and multiple agent runtimes in one engineering workspace.
For people already trusting coding agents with real repositories — and teams that need the same workflows with shared policy, cost visibility and reproducible evidence.
The optimization target: hold the model and task constant, then reduce context replay, wasted turns, retries and brittle tool behavior without trading away verified correctness.
As coding moves from autocomplete to multi-step repository work, orchestration cost and reliability compound. Every unnecessary turn becomes recurring inference spend.
The Apache-2.0 desktop harness is the adoption wedge. The planned paid layer is shared policy, usage governance, managed execution, audit and administration for teams.
The advantage is accumulated harness behavior: runtime-neutral execution, benchmark-driven controller design, durable project state and telemetry — not a clever system prompt.
Projects, source control, real terminal sessions, durable agent threads, verification evidence, usage accounting, and provider-independent runtime control in one desktop workspace.
The same capable model can behave very differently depending on the environment wrapped around it. Trebell focuses on the part a model vendor does not solve for you: the engineering system that turns inference into dependable repository work.
Dumping more of a repository into every turn burns tokens, damages cache locality, and often gives the model less signal rather than more.
Files, terminals, Git, worktrees, credentials, processes, permissions, side effects, crashes, and provider differences all become part of the agent’s actual runtime.
An agent saying “done” is not the same as the software being correct. Real engineering needs evidence, not confidence theater.
Trebell owns the durable environment around the model. Runtime adapters can change; the engineering substrate stays coherent.
Understand what matters, keep the rest retrievable, and preserve stable prompt prefixes where possible.
Deterministic engineering capabilities with permissions, process isolation, secrets, source control, and side-effect control.
Durable threads, checkpoints, crash recovery, usage telemetry, evidence capture, and verification before completion.
Bigger prompts are not automatically better prompts. Trebell treats context as an engineered resource: expensive enough to measure, stable enough to cache, and selective enough to preserve signal.
Trebell measures the harness where it can actually control the result: provider input, cache behavior, model turns, tool calls, failures, retries, and unnecessary inference. No pretending noisy upstream provider latency is a harness breakthrough.
Public Terminal-Bench 4 tasks, pinned model and reasoning effort, comparable permissions, and an independent verifier. Fresh unseen results and causal diagnostics are labeled separately; setup-invalid runs are excluded rather than quietly turned into wins.
Trebell Native matched Codex API's perfect verifier result while using 61.1% less uncached input, 23.1% less output, and 68.3% less agent execution time.
Same task · same model · diagnostic causal evidenceAll three harnesses missed the same core CTR threshold. Native cost $0.0646, versus $0.6973 for Codex API and $0.3319 API-equivalent for OAuth on the sealed comparison.
Shared model failure in this comparison · not counted as a correctness winOn a fresh unseen task, Native matched Codex OAuth's 18 / 19 partial verifier result and exceeded Codex API's 16 / 19. Against API, Native also used 77.9% less raw input and 44.3% less agent time.
Fresh unseen evidence · same remaining miss as OAuthTrebell starts where developers already feel the pain: expensive, brittle coding-agent workflows on real repositories. The long-term opportunity is the neutral control layer teams use across models, runtimes and environments.
Trebell sits beneath model choice. As teams adopt more agents and providers, context, tool execution, verification, permissions and usage governance become shared infrastructure rather than vendor-specific UI.
The planned commercial layer is team policy, shared usage and cost controls, managed execution environments, auditability and administration — capabilities that become more valuable as agent usage scales across an organization.
The current product includes a native agent loop, external runtime adapters, durable SQLite state, Git/worktrees, terminals, browser verification, permission policy, recovery, packaging and a benchmark harness with external verifiers.
Terminal-Bench 4.0 is the current external program. SWE-Bench Pro V2 is the next software-engineering target. The goal is not one pretty benchmark — it is repeatable harness gains across independent tasks.
Codex, Claude Code, Cursor and other agents remain useful runtimes. Trebell's position is different: keep project identity and engineering infrastructure portable across them, while owning the context, policy, verification and telemetry layer the user depends on.
Distribution starts with the downloadable desktop harness, GitHub, reproducible benchmark evidence and developer word of mouth. Team features become the expansion path once multiple people need shared policy, usage controls and managed execution.
Trebell grew out of repeatedly running, debugging and benchmarking coding agents across providers and runtimes. The product is the infrastructure that work kept demanding: better context discipline, reliable tools, recovery, verification and honest efficiency telemetry.
Trebell Code is building the provider-independent engineering layer for coding agents that need to work on real software — with less waste, visible failures and evidence before “done.”