Agent Optimization
Teams change instructions, swap models and adjust settings every week, then have no way to say whether any of it worked. Optimization starts with a measurement.
Every layer of an agent can be tuned and every tuning can be scored. We build the measurement first, establish a baseline, then improve the context, the harness and the routing against it.
Baseline first
Nothing is tuned until the current configuration has been measured.
Gated
No candidate is adopted if it breaks work the system already handles.
Reusable
The measurement stays useful long after the engagement has closed.
Baseline
what you run today
Context
what reaches the model
Harness
the loop around it
Token and cost
routing, right-sizing
Scored run
quality, cost, latency
Held-out gate
adopt or roll back
Fit
Who This Is For
Teams with an agent already in production, where quality varies, cost climbs, and nobody can say which change caused what.
Production agents
Systems live with users where quality now varies run to run
Rising cost
Teams whose token and compute spend grew faster than the value returned
Plateaued quality
Products where tuning stopped producing any visible gain
Large repositories
Codebases exceeding what a single context window can hold
Process
How It Works
Build the measurement
Tasks drawn from your own workload and the failures your team keeps hitting, turned into a scored set you can rerun indefinitely
Establish a baseline
Current configuration measured across the set, with cost and latency recorded alongside quality so tradeoffs stay visible
Optimize the layers
Context and retrieval, the harness around them, and which model handles which role, each tuned against the baseline rather than against opinion
Gate and adopt
Candidates pass a held-out gate before you take them, and a rollback path is written before anything reaches your users
Deliverables
What You Receive
Evaluation set
Scored tasks built from your workload, delivered to your team to keep and rerun after we leave
Optimized configuration
Context, retrieval and harness tuned against the baseline, with the runs behind each decision
Routing plan
Which model handles which role, chosen on measured outcome rather than on list price
Cost position
Token and compute spend measured before and after, with the quality tradeoff stated plainly
In practice
What an Engagement Looks Like
Three shapes this work usually takes. Scope is agreed before anything starts, and none of these depend on you running a particular vendor.
Benchmark only
The measurement built, the baseline recorded and the report handed over. Useful on its own whether or not any tuning follows.
Measure then tune
The usual shape. A baseline first, then candidates compared against it and gated before anything reaches your users.
Spend under a floor
Routing and right-sizing worked against an agreed quality floor, with cost recorded before and after the change.
Pricing
Engagement Options
Buy one column or combine them. Most engagements start with Context, because that is usually where both the quality and the cost problem live.
Most requested
Context
What reaches the model
Typically 3 to 5 weeks
Context and retrieval
The largest share of quality and cost sits in what gets assembled for each call.
- Context assembly reviewed and tuned
- Retrieval quality improvement
- Instruction and skill optimization
- Compaction and window management
- Measured against a baseline
- Held-out gate before adoption
Harness
The loop around the model
Typically 4 to 6 weeks
Harness and tooling
Tools, permissions, checks and control flow, optimized as one executable artifact.
- Harness measured across variants
- Tool surface and permissions tuned
- Checks and control flow
- Recursive routing for large repositories
- Evaluation set you keep
- Documented rollback path
Token & Cost
Spend under a quality floor
Typically 3 to 5 weeks
Routing and right-sizing
Most teams overpay a frontier model for work a smaller one handles at the same quality.
- Token and spend analysis
- Model routing per role
- Local and open model evaluation
- Quantization where accuracy allows
- Quality floor agreed and enforced
- Cost measured before and after
Prices exclude VAT where applicable.
People
Who You Work With
Work is delivered by the Superagentic AI team, supported by consultants and researchers brought in on demand when an engagement calls for depth in a specific area. The same lead stays with you from the first call through to handover, so context does not get rebuilt between phases.
Open source
Projects Behind This Work
Optimization is the work we have published most. You are not buying these tools, and we tune whatever you already run.
MetaHarness
Optimizes the executable harness around a coding agent rather than a prompt string.
DSPy Code
CLI for building and optimizing DSPy programs with GEPA and related optimizers.
RLM Code
Recursive language models over codebases too large to fit a context window.
CodexOpt
Optimizes AGENTS.md and skills with GEPA, discovering targets and applying changes.
TurboAgents
Quantization and right-sizing for agents carrying a token and compute problem.
SuperQode
Harness benchmarking with held-out gates, so an improvement can be proven rather than claimed.
Questions
Frequently Asked Questions
Ready to improve an agent in production
Tell us what you are running today. The measurement stays with your team whether or not you take the tuning.
Working from London and San Francisco. Remote by default, on site where it helps.
