Open-source desktop coding harness v1.3.6 shipped

One desktop harness for AI coding agents.Same model. Less waste. More verified work.

Trebell Code is a provider-independent desktop coding harness between AI models and your repository. Run Trebell Native, Codex, Claude Code, OpenCode and more in one workspace while Trebell owns the context, tools, Git, terminals, state, recovery and verification around them. The goal is simple: reach verified results with less wasted inference.

Shipped
v1.3.6 desktop product
Apache-2.0
Open source
1,279 / 1,279
latest exact test suite
Terminal-Bench 4
external harness evidence

Models keep getting better.
The harness is the durable layer.

A coding model does not ship software by itself. The harness decides what context it sees, which tools it can use, how failures recover, what gets verified, and how much inference is wasted on the way. Trebell is a neutral engineering layer around increasingly capable models.

PRODUCT

A desktop coding harness

Projects, repository intelligence, terminals, Git/worktrees, browser verification, durable threads, recovery and multiple agent runtimes in one engineering workspace.

CUSTOMER

Developers now. Teams next.

For people already trusting coding agents with real repositories — and teams that need the same workflows with shared policy, cost visibility and reproducible evidence.

FOCUS

Improve the harness, not the model

The optimization target: hold the model and task constant, then reduce context replay, wasted turns, retries and brittle tool behavior without trading away verified correctness.

WHY NOW

Agents are becoming long-running workers

As coding moves from autocomplete to multi-step repository work, orchestration cost and reliability compound. Every unnecessary turn becomes recurring inference spend.

BUSINESS MODEL

Open-source adoption → team control plane

The Apache-2.0 desktop harness is the adoption wedge. The planned paid layer is shared policy, usage governance, managed execution, audit and administration for teams.

WHY IT COMPOUNDS

Controller + substrate + evidence

The advantage is accumulated harness behavior: runtime-neutral execution, benchmark-driven controller design, durable project state and telemetry — not a clever system prompt.

StageShipped v1.3.6
Quality gate1,279 / 1,279
BenchmarkingTerminal-Bench 4.0
BuilderSolo technical founder

A real engineering workspace.
Not a chat box with a terminal bolted on.

Projects, source control, real terminal sessions, durable agent threads, verification evidence, usage accounting, and provider-independent runtime control in one desktop workspace.

Trebell Code desktop workspace in dark mode Trebell Code desktop workspace in light mode
Actual Trebell Code interface captured from the current desktop build.

AI coding is becoming infrastructure.
The harness is the bottleneck.

The same capable model can behave very differently depending on the environment wrapped around it. Trebell focuses on the part a model vendor does not solve for you: the engineering system that turns inference into dependable repository work.

⌘01 / 03

Context becomes a tax.

Dumping more of a repository into every turn burns tokens, damages cache locality, and often gives the model less signal rather than more.

  • IN TREBELL Repository intelligence is retrievable instead of eagerly stuffed into every prompt.
  • IN TREBELL Large outputs can be virtualized while keeping hot previews available to the model.
◇02 / 03

Agents operate in brittle worlds.

Files, terminals, Git, worktrees, credentials, processes, permissions, side effects, crashes, and provider differences all become part of the agent’s actual runtime.

  • IN TREBELL Deterministic tools sit behind shared policy instead of bespoke provider behavior.
  • IN TREBELL State, source control, process handling, and recovery belong to the harness.
✓03 / 03

Completion is not proof.

An agent saying “done” is not the same as the software being correct. Real engineering needs evidence, not confidence theater.

  • IN TREBELL Verification evidence is part of the execution loop rather than an optional afterthought.
  • IN TREBELL Repairs are measured end-to-end instead of celebrating isolated micro-metrics.

One engineering control plane.
Multiple capable agents.

Trebell owns the durable environment around the model. Runtime adapters can change; the engineering substrate stays coherent.

01 ⌘

Repository & Context Intelligence

Understand what matters, keep the rest retrievable, and preserve stable prompt prefixes where possible.

repo graphretrievalvirtualizationcache-aware
02 ◇

Tools, Policy & Execution

Deterministic engineering capabilities with permissions, process isolation, secrets, source control, and side-effect control.

terminalfilesgitpolicyMCP
03 ✓

State, Trace & Verification

Durable threads, checkpoints, crash recovery, usage telemetry, evidence capture, and verification before completion.

statetracerecoveryevidence
Runtime adapters
Trebell NativeCodexClaude CodeOpenCodeACP+ future agents

Give the model what matters.
Keep the rest retrievable.

Bigger prompts are not automatically better prompts. Trebell treats context as an engineered resource: expensive enough to measure, stable enough to cache, and selective enough to preserve signal.

  • Progressive discoveryExpose specialized tools and MCP capability only when the task needs them.
  • Large-output virtualizationKeep bulky command output retrievable instead of replaying it through every turn.
  • Stable identityKeep project/thread state independent of whichever model provider is active.

Optimize useful coding work.
Not vanity metrics.

Trebell measures the harness where it can actually control the result: provider input, cache behavior, model turns, tool calls, failures, retries, and unnecessary inference. No pretending noisy upstream provider latency is a harness breakthrough.

  • Token & cache telemetrySee where context is paid for, reused, or wasted.
  • Tool/retry accountingMeasure how much orchestration the model needed to reach a verified result.
  • Regression disciplineAn optimization only counts when the end-to-end task gets better.

Same model. Same task.
Measure the harness.

Public Terminal-Bench 4 tasks, pinned model and reasoning effort, comparable permissions, and an independent verifier. Fresh unseen results and causal diagnostics are labeled separately; setup-invalid runs are excluded rather than quietly turned into wins.

Shadow-relay · controlled causal rerun 66.3% less input vs API
8 / 8 = 8 / 8

Same result, much less traffic

Trebell Native matched Codex API's perfect verifier result while using 61.1% less uncached input, 23.1% less output, and 68.3% less agent execution time.

Same task · same model · diagnostic causal evidence
CTR optimization · shared model miss 90.7% lower cost vs API
3 / 4 = 3 / 4

Same failure. A fraction of the spend.

All three harnesses missed the same core CTR threshold. Native cost $0.0646, versus $0.6973 for Codex API and $0.3319 API-equivalent for OAuth on the sealed comparison.

Shared model failure in this comparison · not counted as a correctness win
Legacy utility triage · fresh unseen task 77.9% less input vs API
18 / 19 = 18 / 19

Matched OAuth. Beat API quality.

On a fresh unseen task, Native matched Codex OAuth's 18 / 19 partial verifier result and exceeded Codex API's 16 / 19. Against API, Native also used 77.9% less raw input and 44.3% less agent time.

Fresh unseen evidence · same remaining miss as OAuth
BENCHMARKTerminal-Bench 4.0
MODELSame model per comparison
REASONINGPinned equally
SCORINGIndependent verifier

Start with individual developers.
Expand into the team layer.

Trebell starts where developers already feel the pain: expensive, brittle coding-agent workflows on real repositories. The long-term opportunity is the neutral control layer teams use across models, runtimes and environments.

Market

More agent usage means more infrastructure to manage.

Trebell sits beneath model choice. As teams adopt more agents and providers, context, tool execution, verification, permissions and usage governance become shared infrastructure rather than vendor-specific UI.

Business model

Open-source developer adoption. Paid team infrastructure.

The planned commercial layer is team policy, shared usage and cost controls, managed execution environments, auditability and administration — capabilities that become more valuable as agent usage scales across an organization.

Execution

Deep product work, not a landing-page prototype.

The current product includes a native agent loop, external runtime adapters, durable SQLite state, Git/worktrees, terminals, browser verification, permission policy, recovery, packaging and a benchmark harness with external verifiers.

Next milestone

Prove the pattern across more unseen software work.

Terminal-Bench 4.0 is the current external program. SWE-Bench Pro V2 is the next software-engineering target. The goal is not one pretty benchmark — it is repeatable harness gains across independent tasks.

Competition

Vendor-native agents optimize one stack. Trebell stays neutral.

Codex, Claude Code, Cursor and other agents remain useful runtimes. Trebell's position is different: keep project identity and engineering infrastructure portable across them, while owning the context, policy, verification and telemetry layer the user depends on.

Go to market

Win developers with a useful open-source product first.

Distribution starts with the downloadable desktop harness, GitHub, reproducible benchmark evidence and developer word of mouth. Team features become the expansion path once multiple people need shared policy, usage controls and managed execution.

Founder-market fit

Built end-to-end by a solo technical founder who is already living the problem.

Trebell grew out of repeatedly running, debugging and benchmarking coding agents across providers and runtimes. The product is the infrastructure that work kept demanding: better context discipline, reliable tools, recovery, verification and honest efficiency telemetry.

The model will keep changing.
The engineering layer should compound.

Trebell Code is building the provider-independent engineering layer for coding agents that need to work on real software — with less waste, visible failures and evidence before “done.”