Context pruner
Decide, item by item, whether an agent's context item still matters for the current task, and drop the rest unchanged.
Decide, item by item, whether an agent's context item still matters for the current task, and drop the rest unchanged. Pattern: confidence_routing. Model: jev-1.13.0. Outage rule: review.
When to use it
- An agent's context fills with old tool calls, tool results and messages, and you pay for every token on every step.
- You want to keep or drop whole items. Nothing gets reworded, so nothing gets distorted.
When not to use it
- You want a summary of the dropped items. That is text generation; keep an LLM for it.
- Items are tiny and few. The saving will not cover the call.
- The rule is mechanical, such as dropping tool results older than a set number of steps. Keep that in code.
Questions
| Id | Type | Asks | Answers | Gating |
|---|---|---|---|---|
still_matters | noul | An agent is working on task. Would the agent need item in its context to finish that task correctly? | true or false | yes |
Reading the result
- One run per item. Drop an item only when
overallActionisautoandrouteisdrop. Every other outcome keeps it. - The thresholds lean toward keeping:
falseAtis low, so the model must be sure an item no longer matters before it goes. - Pass
options.metadata.tokensBeforeandtokensAfteron the run to book context tokens saved. - Keep the latest user message and the system prompt out of the candidates. Pin them in code.
Spec
Save it as bandwise/sets/context-pruner.json and edit it for your data.
{
"schemaVersion": 1,
"model": "jev-1.13.0",
"input": {
"schema": {
"type": "object",
"required": [
"task",
"item"
],
"properties": {
"task": {
"type": "string",
"maxLength": 2000
},
"item": {
"type": "object",
"required": [
"kind",
"content"
],
"properties": {
"kind": {
"type": "string",
"enum": [
"tool_call",
"tool_result",
"message"
]
},
"content": {
"type": "string"
}
}
}
}
}
},
"stages": [
{
"id": "prune",
"questions": {
"still_matters": {
"type": "noul",
"instructions": "An agent is working on `task`. Would the agent need `item` in its context to finish that task correctly?",
"criteria": {
"true": "The item holds facts, decisions, file contents, errors or instructions the agent still needs for the task.",
"false": "The item is finished business: superseded output, a dead end already abandoned, or chatter with nothing the task depends on."
},
"meta": {
"label": "Still matters"
}
}
}
}
],
"policies": {
"still_matters": {
"type": "noul",
"gating": true,
"noul": {
"trueAt": 0.7,
"falseAt": 0.15,
"reviewMargin": 0.1
},
"actions": {
"high": {
"kind": "auto"
},
"medium": {
"kind": "fallback",
"config": {
"kind": "value",
"value": true
}
},
"low": {
"kind": "fallback",
"config": {
"kind": "value",
"value": true
}
}
}
}
},
"routes": [
{
"when": {
"q": "still_matters",
"eq": false
},
"output": "drop"
}
],
"defaultRoute": "keep",
"savings": {
"comparatorModel": "claude-haiku-4-5",
"estOutputTokensPerQuestion": 20,
"kind": "context_pruned"
},
"onUnavailable": "review"
}Example states
The expected outcome is what a person would decide. It is not a recorded model answer.
Current test failure output
Expected: keep: it is the error the agent is fixing.
{
"task": "Fix the failing test in src/billing/invoice.test.ts without changing the public API.",
"item": {
"kind": "tool_result",
"content": "FAIL src/billing/invoice.test.ts > totals include tax. Expected 118.00, received 100.00."
}
}Directory listing from an unrelated package
Expected: drop: the agent looked there and moved on.
{
"task": "Fix the failing test in src/billing/invoice.test.ts without changing the public API.",
"item": {
"kind": "tool_result",
"content": "packages/marketing-site: README.md, next.config.js, src/, public/"
}
}Borderline cases
One case near the line for each question. Use them to test your wording before you trust the thresholds.
still_matters
An earlier version of the file the agent has since edited. Mostly superseded, but it shows what the public API looked like.
{
"task": "Fix the failing test in src/billing/invoice.test.ts without changing the public API.",
"item": {
"kind": "tool_result",
"content": "src/billing/invoice.ts (before edits): export function total(lines) { return sum(lines); }"
}
}Try it
Save an example state as state.json, then run the spec locally. Local mode makes no network call and needs no key; answers are synthetic unless a recorded fixture matches.
pnpm bandwise run --local bandwise/sets/context-pruner.json state.jsonBandwise is an independent product built on TypeSafe's System One models. It is not TypeSafe's documentation. For the System One models themselves, see docs.typesafe.ai.