The Gemini Interactions API is Google’s recommended surface for new agent projects: one call shape covering model thoughts, tool steps, and stateful follow-ups. This tutorial builds thinking-router.ts on that API, a TypeScript program that classifies each coding task, selects a model-supported thinking level, and records latency, token usage, and estimated cost. The router adds its own overhead, and it does not guarantee savings in either time or money.
The Hidden Cost of a Static Thinking Setup
Consider a hypothetical autonomous coding agent that spends eight seconds and thousands of thinking tokens to fix a one-character typo. The scenario is invented. It is not a measured trace, and its numbers do not describe observed Gemini behavior. It illustrates a failure mode that is easy to build into an agent when the agent treats Gemini thinking budgets as one global setting rather than a per-task decision.
One premise needs correcting first. Gemini already adjusts thinking dynamically when the caller supplies no control. The stable gemini-3.8-flash model supports low, medium, and high, and defaults to medium (per the current Gemini thinking documentation). A routing layer is therefore not the only way to avoid heavy reasoning on trivial work. Its value is that it makes the choice explicit and measurable. That choice then has to beat the model’s own default across complete coding tasks. Comparing high and low on isolated prompts does not answer that question.
The Gemini Interactions API is Google’s recommended surface for new agent projects: one call shape covering model thoughts, tool steps, and stateful follow-ups. This tutorial builds thinking-router.ts on that API, a TypeScript program that classifies each coding task, selects a model-supported thinking level, and records latency, token usage, and estimated cost. The router adds its own overhead, and it does not guarantee savings in either time or money.
Thinking Levels Across the Two Gemini APIs
Gemini 3 Thinking Levels Versus Gemini 2.5 Thinking Budgets
The Gemini API offers two request surfaces. Do not mix them.
- Interactions API. Google recommends this generally available API for new agent projects. Set thinking with
generation_config: { thinking_level: "low" }. generateContent. Google still fully supports this API, and projects that need Batch API or explicit caching still require it because Interactions does not yet offer those features. Set thinking withconfig: { thinkingConfig: { thinkingLevel: "low" } }.
Neither documented request shape includes the outline’s reasoningEffort.
Two companion posts cover the generateContent side of this space: Optimizing Gemini 3.8 Flash for Autonomous Coding Agents builds a thinking-level agent harness with tool retries, and Gemini 3.8 Flash Reasoning Effort: Balancing Latency and Costs builds latency-driven tier cascades. This article focuses on the Interactions API surface instead.
On Gemini 3 models, the control is a qualitative thinkingLevel. For gemini-3.8-flash, the valid values are low, medium (the default), and high. Passing minimal produces an error.
The numeric thinkingBudget applies to Gemini 2.5, for example config: { thinkingConfig: { thinkingBudget: 1024 } }. The valid ranges differ by model:
gemini-2.5-flashaccepts 0 to 24,576.gemini-2.5-proaccepts 128 to 32,768 and cannot turn thinking off.- A value of
-1requests dynamic thinking.
A numeric budget guides the model rather than capping it; Google has documented that actual usage can fall below or exceed the requested value. These ranges are the historical Gemini 2.5 values – the current thinking documentation describes level-based control for the Gemini 3 series.
Google bills thinking tokens as output tokens, and the charge covers the model’s full generated thoughts, not only the optional thought summaries. For paid standard-tier gemini-3.8-flash, pricing as of September 28, 2026 is $0.75 per million input tokens and $3.75 per million output tokens, with thinking tokens counted as output. Both rates are scheduled to double on January 1, 2027 (per the official Gemini pricing page, checked September 28, 2026). Extra reasoning can also delay the first non-thinking output token.
Why Unexamined Budgets Break Agent Loops
An agent issues many calls per task: planning, editing, repairing, and verifying. Each call’s thinking cost and latency add up, so what matters is cumulative billed tokens and elapsed time across the loop, not the setting on any single request. In CI/CD pipelines or interactive developer tools, that latency accumulates at every step.
Calling an unspecified setting „maximum” is wrong. Gemini documents model defaults and dynamic behavior, not a universal maximum-budget default.
A thinking setting is also not a hard per-request spending cap. The output-token cutoff is a separate control with its own risk. If the caller sets it too low, generation can stop partway through thinking. The caller gets truncated or empty visible output and still pays for the thinking tokens consumed before the cutoff.
The Core Problem: Runaway Token Costs on Simple Tasks
Take a „rename variable” instruction. A fair evaluation holds the input, model, tools, and success criterion constant and compares three configurations: the default, low, and high.
Whether a short instruction describes a simple job depends on the work needed to produce and verify a safe patch. Renaming a local variable in one function is trivial. Renaming an exported symbol used across a package is not, even though both instructions fit on one line.
Two claims in the outline are unverified:
- No trace establishes that a Gemini rename consumed „thousands” of thought tokens or took eight seconds.
- No evidence shows that a static
highsetting performs worse than Gemini’s default for any particular task distribution.
Both are hypotheses to test, not results. „Runaway” also should not suggest that a numeric budget enforces an exact ceiling, since 2.5 models can exceed their requested budget.
The operational risk extends beyond a single expensive call. Gemini evaluates spend-based rate limits over rolling ten-minute windows, and Google lists $10, $50, and $200 per ten minutes for paid tiers 1, 2, and 3. Exceeding the limit returns 429 RESOURCE_EXHAUSTED (per the Gemini rate-limit documentation). An agent fleet that applies heavy thinking to every task reaches that limit sooner.
A one-size-fits-all setting fails because task difficulty varies. Complexity-based routing responds to that variation, but only if you count its full cost. Classifier overhead, retries, tool calls, and verification calls all belong in the total. A router can save thought tokens on the coding request and still lose money or time overall.
A router can save thought tokens on the coding request and still lose money or time overall.
Designing a Dynamic Thinking Router
Defining a Task Complexity Taxonomy
The router sorts tasks into four tiers:
- Trivial edits, such as renames and import fixes.
- Localized bug fixes confined to one file or one function.
- Multi-file refactors that coordinate changes across modules.
- Architecture-level synthesis: design work, API boundaries, and dependency changes.
Classification must rely only on information available before the coding call:
- intended operation
- affected-file count
- diff size, if a diff already exists
- dependency or public API changes
- test scope
- missing metadata
If the agent must generate a diff first, diff size is not a free input to the classifier.
The router uses a deterministic first-pass classifier instead of a second LLM call. That avoids paying for an extra request and avoids sending repository context to one. A model-assisted classifier earns a place only if a benchmark run shows it raises correctness or lowers total task cost compared with the deterministic version, with its own latency, tokens, failures, and price counted in the task totals.
Mapping Complexity Tiers to Thinking Levels
The outline proposed numeric budget ranges with min/max clamps. That design fits Gemini 2.5, not gemini-3.8-flash. For Gemini 3, the correct guardrail is an allowlist of model-supported levels, keyed by model ID, that rejects unsupported values before the router makes any API call.
The starting map for gemini-3.8-flash:
| Tier | Level |
|---|---|
| Trivial edits | low |
| Localized fixes | low or medium |
| Multi-file refactors | medium |
| Architecture synthesis | high |
These are proposed settings for evaluation. Google does not prescribe these thresholds, and this tutorial has not validated them for correctness. Treat them as hypotheses until benchmark runs produce correctness and cost results.
Three fallback rules make the router conservative:
- Low-confidence classifications route to
medium, the documented default. - Any single missing metadata field, including an undeclared public-API or dependency flag, drops confidence below the floor and triggers that fallback.
- If the router flags an edit as a public-API or dependency change, that edit never runs at
low. The router models no other risk categories.
A route decision is a prediction to validate against tests. It never authorizes skipping tests or human approval.
A route decision is a prediction to validate against tests. It never authorizes skipping tests or human approval.
Some teams keep a Gemini 2.5 Flash integration. In a separately labeled compatibility path, the router can validate a numeric budget against 0 to 24,576 or -1. Never apply that clamp to a Gemini 3 request.
Building the Harness: thinking-router.ts
Project Setup and Dependencies
@google/genai is the maintained JavaScript SDK. Its official repository lists v2.24.0 as the September 22, 2026 release.
Node version matters here. SDK 2.x accepts Node 20+, but Node 20 reached end-of-life in April 2026 (per the Node.js release schedule). This tutorial targets Node 24 LTS and pins the SDK to exactly 2.24.0, the current latest @google/genai release.
Install the dependencies:
npm init -y
npm install --save-exact @google/genai@2.24.0
npm install zod
npm install --save-dev typescript @types/node--save-exact records @google/genai as exactly 2.24.0, so a later 2.x minor release cannot silently change the typings the example compiles against. Commit the generated package-lock.json to record the exact dependency versions used for the tested example. Merge the following fields into the package.json that npm created; do not replace its existing dependency sections.
{
"name": "thinking-router",
"version": "0.1.0",
"private": true,
"type": "module",
"engines": { "node": ">=24" },
"scripts": {
"build": "tsc",
"typecheck": "tsc --noEmit",
"start": "node dist/thinking-router.js",
"test": "npm run build && node --test \"dist/**/*.test.js\""
}
}Run npm run typecheck as a gate after installing. It must exit with code 0, which confirms that the request and response fields used below match the pinned SDK typings.
Add a tsconfig.json:
{
"compilerOptions": {
"target": "ES2022",
"module": "NodeNext",
"moduleResolution": "NodeNext",
"strict": true,
"outDir": "dist",
"rootDir": "src",
"skipLibCheck": true
},
"include": ("src")
}The router passes the key explicitly with new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY }). Never ship the API key in browser-distributed code.
Defining the Task Classification Schema with Zod
The field names and tier thresholds below are design choices for this tutorial. They are not Gemini API fields. The schema guarantees that every classification the router receives is structurally valid, whether it comes from the heuristic or from a future LLM classifier.
import { z } from "zod";
const TaskTypeSchema = z.enum((
"trivial_edit",
"localized_fix",
"multi_file_refactor",
"architecture_synthesis",
));
type TaskType = z.infer<typeof TaskTypeSchema>;
const TaskComplexitySchema = z.object({
taskType: TaskTypeSchema,
complexityScore: z.number().int().min(0).max(100),
confidence: z.number().min(0).max(1),
highRisk: z.boolean(),
fallbackReason: z.string().optional(),
});
type TaskComplexity = z.infer<typeof TaskComplexitySchema>;The same pattern applies to a model-assisted classifier. Google’s Interactions example sends a JSON Schema through response_format: { type: "text", mime_type: "application/json", schema }, then parses JSON.parse(interaction.output_text) with Zod. Parse and validate a schema-shaped response before use anyway.
Classifying Incoming Coding Tasks
The classifier scores a task from its operation type, affected-file count, any pre-existing diff, API and dependency changes, and test scope. Each missing field, including an undeclared changesPublicApi or changesDependencies flag, lowers confidence enough on its own to fall below the routing floor. A negative diff size cannot lower the score. If validation fails, the function returns a conservative classification instead of throwing.
interface CodingTask {
id: string;
operation:
| "rename"
| "fix_import"
| "bug_fix"
| "refactor"
| "feature"
| "design"
| "unknown";
instruction: string;
affectedFiles?: number;
existingDiffLines?: number;
changesPublicApi?: boolean;
changesDependencies?: boolean;
testScope?: "none" | "unit" | "integration" | "full";
}
const OPERATION_BASE_SCORE: Record<CodingTask("operation"), number> = {
rename: 5,
fix_import: 5,
bug_fix: 25,
refactor: 45,
feature: 55,
design: 75,
unknown: 40,
};
export const CONSERVATIVE_FALLBACK: TaskComplexity = TaskComplexitySchema.parse({
taskType: "multi_file_refactor",
complexityScore: 50,
confidence: 0,
highRisk: true,
fallbackReason: "classification failed validation",
});
export function classifyTask(task: CodingTask): TaskComplexity {
const missing: string() = ();
if (task.affectedFiles === undefined) missing.push("affectedFiles");
if (task.testScope === undefined) missing.push("testScope");
if (task.changesPublicApi === undefined) missing.push("changesPublicApi");
if (task.changesDependencies === undefined) missing.push("changesDependencies");
if (task.operation === "unknown") missing.push("operation");
const files = Math.max(task.affectedFiles ?? 1, 1);
let score = OPERATION_BASE_SCORE(task.operation);
score += Math.min(files - 1, 10) * 3;
if (task.existingDiffLines !== undefined) {
score += Math.max(0, Math.min(Math.floor(task.existingDiffLines / 50), 10));
}
if (task.changesPublicApi) score += 10;
if (task.changesDependencies) score += 10;
if (task.testScope === "integration") score += 5;
if (task.testScope === "full") score += 10;
score = Math.min(100, Math.max(0, score));
const taskType: TaskType =
score < 15
? "trivial_edit"
: score < 40
? "localized_fix"
: score < 70
? "multi_file_refactor"
: "architecture_synthesis";
const candidate = {
taskType,
complexityScore: score,
confidence: Math.max(0.2, 0.9 - missing.length * 0.35),
highRisk: Boolean(task.changesPublicApi || task.changesDependencies),
fallbackReason:
missing.length > 0 ? `missing metadata: ${missing.join(", ")}` : undefined,
};
const parsed = TaskComplexitySchema.safeParse(candidate);
return parsed.success ? parsed.data : CONSERVATIVE_FALLBACK;
}These rules reflect the rename distinction from earlier. A local rename with its API and dependency flags declared false scores 5 and classifies as trivial. A rename that changes a public API scores 15, becomes a localized fix, and gets flagged high-risk. A rename whose API or dependency flags are omitted loses enough confidence to route to medium rather than low.
Mapping Complexity to Thinking Levels
The outline’s resolveThinkingBudget() becomes resolveThinkingLevel(). Mapping a tier to a numeric thinkingBudget is a 2.5-style implementation. Because the allowlist is keyed by model ID, adding a model forces you to declare the levels it accepts.
const SUPPORTED_LEVELS = {
"gemini-3.8-flash": ("low", "medium", "high"),
} as const;
type SupportedModel = keyof typeof SUPPORTED_LEVELS;
type ThinkingLevel = (typeof SUPPORTED_LEVELS)(SupportedModel)(number);
const TIER_LEVELS: Record<TaskType, ThinkingLevel> = {
trivial_edit: "low",
localized_fix: "medium",
multi_file_refactor: "medium",
architecture_synthesis: "high",
};
const CONFIDENCE_FLOOR = 0.6;
export function resolveThinkingLevel(
model: SupportedModel,
c: TaskComplexity,
): ThinkingLevel {
let level: ThinkingLevel = TIER_LEVELS(c.taskType);
if (c.taskType === "localized_fix" && c.complexityScore < 30 && !c.highRisk) {
level = "low";
}
if (c.confidence < CONFIDENCE_FLOOR) level = "medium";
if (c.highRisk && level === "low") level = "medium";
const allowed: readonly string() = SUPPORTED_LEVELS(model);
if (!allowed.includes(level)) {
throw new Error(`Thinking level "${level}" is not supported by ${model}`);
}
return level;
}Injecting Parameters into the Gemini API Call
The example assumes the documented Interactions request and response shape for the pinned @google/genai version, and npm run typecheck confirms that these fields compile against the SDK typings (see the Interactions API reference). The router reads response text from interaction.output_text and usage from interaction.usage. It records each usage counter as a number only when the response supplies a finite number. It records a missing or renamed field as null rather than 0, so a field-name mismatch surfaces as unknown usage instead of a false zero.
Interactions defaults to store: true, which retains interactions for 55 days on the paid tier and one day on the free tier (per the Interactions API data-retention documentation). Coding prompts often contain proprietary source, so this example requests store: false, which the API describes as disabling stored interactions. The trade-off: store: false also disables continuation through previous_interaction_id.
import { GoogleGenAI } from "@google/genai";
interface RoutePlan {
classification: TaskComplexity;
level: ThinkingLevel;
}
interface UsageSnapshot {
inputTokens: number | null;
outputTokens: number | null;
thoughtTokens: number | null;
cachedTokens: number | null;
toolUseTokens: number | null;
}
interface AgentTaskResult {
plan: RoutePlan;
outputText: string | undefined;
usage: UsageSnapshot | null;
}
const REQUEST_TIMEOUT_MS = 120_000;
function toCount(v: unknown): number | null {
return typeof v === "number" && Number.isFinite(v) ? v : null;
}
function withTimeout<T>(p: Promise<T>, ms: number): Promise<T> {
let timer: ReturnType<typeof setTimeout> | undefined;
const timeout = new Promise<never>((_, reject) => {
timer = setTimeout(() => reject(new Error(`request timed out after ${ms} ms`)), ms);
});
return Promise.race((p, timeout)).finally(() => clearTimeout(timer));
}
class TaskRunError extends Error {
constructor(
message: string,
readonly plan: RoutePlan,
readonly original: unknown,
) {
super(message, { cause: original });
this.name = "TaskRunError";
}
}
export async function runAgentTask(
ai: GoogleGenAI,
model: SupportedModel,
task: CodingTask,
timeoutMs: number = REQUEST_TIMEOUT_MS,
): Promise<AgentTaskResult> {
const classification = classifyTask(task);
const level = resolveThinkingLevel(model, classification);
const plan: RoutePlan = { classification, level };
try {
const interaction = await withTimeout(
ai.interactions.create({
model,
input: task.instruction,
generation_config: { thinking_level: level },
store: false,
}),
timeoutMs,
);
const u = interaction.usage;
return {
plan,
outputText: interaction.output_text,
usage: u
? {
inputTokens: toCount(u.total_input_tokens),
outputTokens: toCount(u.total_output_tokens),
thoughtTokens: toCount(u.total_thought_tokens),
cachedTokens: toCount(u.total_cached_tokens),
toolUseTokens: toCount(u.total_tool_use_tokens),
}
: null,
};
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
throw new TaskRunError(message, plan, err);
}
}Despite its name, runAgentTask() makes a single text-generation request. It does not turn the router into an autonomous agent. Applying edits, running tests, handling tool results, and obtaining approval for consequential operations all require additional application code. Google assigns custom function execution and argument validation to the application.
Warning: The example bounds each call with a wall-clock timeout (REQUEST_TIMEOUT_MS, 120 seconds by default), so a hung request cannot stall the task loop indefinitely. The timeout only stops waiting. It does not cancel the underlying request, which may still complete and incur charges. Production integrations should add request cancellation through a mechanism the SDK supports, and test both limits.
Structured Latency and Token Cost Tracking
TaskMetrics uses only Interactions usage fields. The generateContent fields (promptTokenCount, candidatesTokenCount, thoughtsTokenCount) belong to the other API, and the code deliberately excludes them.
The cost estimate uses the paid standard-tier rates above, valid through December 31, 2026, and applies only to uncached, tool-free calls:
input × 0.75 / 1M + (output + thought) × 3.75 / 1M
This formula treats total_output_tokens as excluding thought tokens, which matches the code’s default OUTPUT_INCLUDES_THOUGHTS = false. If the Interactions API counts thoughts inside total_output_tokens, set the constant to true; otherwise the estimate double counts thinking. The usage sample in the Interactions documentation reports total_input_tokens, total_output_tokens, and total_thought_tokens as disjoint fields that sum to total_tokens, so output excludes thoughts.
In every other case, including any usage counter missing from the response, the router returns null rather than a misleading figure. It records a call that completes without visible output text as empty_output, not ok, because Google still bills that call’s thought tokens. It also records failures, rate limits, and classifier fallbacks, though a failed call has unknown usage and cost when the response carries no usage metadata. The router never logs raw prompts and truncates error messages to 500 characters.
interface TaskMetrics {
taskId: string;
model: SupportedModel;
taskType: TaskType | null;
thinkingLevel: ThinkingLevel | null;
classificationFallback: string | null;
status: "ok" | "empty_output" | "rate_limited" | "error";
latencyMs: number;
outputChars: number | null;
inputTokens: number | null;
outputTokens: number | null;
thoughtTokens: number | null;
cachedTokens: number | null;
toolUseTokens: number | null;
estimatedCostUsd: number | null;
pricingBasis: string;
errorMessage?: string;
}
const INPUT_USD_PER_M = 0.75;
const OUTPUT_USD_PER_M = 3.75;
const PRICING_BASIS = `gemini-3.8-flash paid standard tier, $${INPUT_USD_PER_M}/M input, $${OUTPUT_USD_PER_M}/M output incl. thinking, valid through 2026-12-31`;
const PRICING_VALID_UNTIL = Date.parse("2026-12-31T23:59:59Z");
const OUTPUT_INCLUDES_THOUGHTS = false;
export function estimateCostUsd(
u: UsageSnapshot | null,
now: number = Date.now(),
): number | null {
if (!u) return null;
const { inputTokens, outputTokens, thoughtTokens, cachedTokens, toolUseTokens } = u;
if (inputTokens === null || outputTokens === null || thoughtTokens === null) return null;
if (cachedTokens === null || toolUseTokens === null) return null;
if (cachedTokens > 0 || toolUseTokens > 0) return null;
if (now > PRICING_VALID_UNTIL) return null;
const billedOutput = OUTPUT_INCLUDES_THOUGHTS ? outputTokens : outputTokens + thoughtTokens;
return (
(inputTokens * INPUT_USD_PER_M) / 1_000_000 +
(billedOutput * OUTPUT_USD_PER_M) / 1_000_000
);
}
export function isRateLimited(err: unknown): boolean {
const status = (err as { status?: unknown } | null)?.status;
const message = err instanceof Error ? err.message : String(err);
return status === 429 || status === "429" || message.includes("RESOURCE_EXHAUSTED");
}
function logMetrics(m: TaskMetrics): void {
console.log(JSON.stringify({ ts: new Date().toISOString(), ...m }));
}
export async function trackedRun(
ai: GoogleGenAI,
model: SupportedModel,
task: CodingTask,
timeoutMs: number = REQUEST_TIMEOUT_MS,
): Promise<TaskMetrics> {
const start = performance.now();
try {
const result = await runAgentTask(ai, model, task, timeoutMs);
const metrics: TaskMetrics = {
taskId: task.id,
model,
taskType: result.plan.classification.taskType,
thinkingLevel: result.plan.level,
classificationFallback: result.plan.classification.fallbackReason ?? null,
status: result.outputText?.trim() ? "ok" : "empty_output",
latencyMs: performance.now() - start,
outputChars: result.outputText?.length ?? null,
inputTokens: result.usage?.inputTokens ?? null,
outputTokens: result.usage?.outputTokens ?? null,
thoughtTokens: result.usage?.thoughtTokens ?? null,
cachedTokens: result.usage?.cachedTokens ?? null,
toolUseTokens: result.usage?.toolUseTokens ?? null,
estimatedCostUsd: estimateCostUsd(result.usage),
pricingBasis: PRICING_BASIS,
};
logMetrics(metrics);
return metrics;
} catch (err) {
const plan = err instanceof TaskRunError ? err.plan : null;
const original = err instanceof TaskRunError ? err.original : err;
const metrics: TaskMetrics = {
taskId: task.id,
model,
taskType: plan?.classification.taskType ?? null,
thinkingLevel: plan?.level ?? null,
classificationFallback: plan?.classification.fallbackReason ?? null,
status: isRateLimited(original) ? "rate_limited" : "error",
latencyMs: performance.now() - start,
outputChars: null,
inputTokens: null,
outputTokens: null,
thoughtTokens: null,
cachedTokens: null,
toolUseTokens: null,
estimatedCostUsd: null,
pricingBasis: PRICING_BASIS,
errorMessage: (err instanceof Error ? err.message : String(err)).slice(0, 500),
};
logMetrics(metrics);
return metrics;
}
}The timer starts before classification, so recorded latency includes classifier overhead. The router logs a 429 RESOURCE_EXHAUSTED spend-limit failure as rate_limited, separate from model or task errors. After a rate limit, the entry point stops issuing the remaining tasks instead of sending requests that will likely fail too. The router does not retry; production callers need a documented, bounded retry policy where the API permits retrying.
Full Working Example and Sample Output
The consolidated file below uses one API surface throughout: ai.interactions.create, generation_config.thinking_level, interaction.output_text, and interaction.usage. It routes tasks and calls the model. It works only if the model, API, and SDK details above hold, and it does not apply or verify patches.
Save the file as src/thinking-router.ts, then build and run it on Node 24 LTS. The commands assume bash or zsh; read -rs reads the key without echoing it and keeps it out of shell history. The program writes per-task JSON metrics to stdout and the summary table to stderr, so npm start > metrics.jsonl produces a parseable JSONL file.
read -rs GEMINI_API_KEY && export GEMINI_API_KEY
npm run build
npm startimport { pathToFileURL } from "node:url";
import { Console } from "node:console";
import { GoogleGenAI } from "@google/genai";
import { z } from "zod";
const TaskTypeSchema = z.enum((
"trivial_edit",
"localized_fix",
"multi_file_refactor",
"architecture_synthesis",
));
type TaskType = z.infer<typeof TaskTypeSchema>;
const TaskComplexitySchema = z.object({
taskType: TaskTypeSchema,
complexityScore: z.number().int().min(0).max(100),
confidence: z.number().min(0).max(1),
highRisk: z.boolean(),
fallbackReason: z.string().optional(),
});
type TaskComplexity = z.infer<typeof TaskComplexitySchema>;
interface CodingTask {
id: string;
operation:
| "rename"
| "fix_import"
| "bug_fix"
| "refactor"
| "feature"
| "design"
| "unknown";
instruction: string;
affectedFiles?: number;
existingDiffLines?: number;
changesPublicApi?: boolean;
changesDependencies?: boolean;
testScope?: "none" | "unit" | "integration" | "full";
}
const OPERATION_BASE_SCORE: Record<CodingTask("operation"), number> = {
rename: 5,
fix_import: 5,
bug_fix: 25,
refactor: 45,
feature: 55,
design: 75,
unknown: 40,
};
export const CONSERVATIVE_FALLBACK: TaskComplexity = TaskComplexitySchema.parse({
taskType: "multi_file_refactor",
complexityScore: 50,
confidence: 0,
highRisk: true,
fallbackReason: "classification failed validation",
});
export function classifyTask(task: CodingTask): TaskComplexity {
const missing: string() = ();
if (task.affectedFiles === undefined) missing.push("affectedFiles");
if (task.testScope === undefined) missing.push("testScope");
if (task.changesPublicApi === undefined) missing.push("changesPublicApi");
if (task.changesDependencies === undefined) missing.push("changesDependencies");
if (task.operation === "unknown") missing.push("operation");
const files = Math.max(task.affectedFiles ?? 1, 1);
let score = OPERATION_BASE_SCORE(task.operation);
score += Math.min(files - 1, 10) * 3;
if (task.existingDiffLines !== undefined) {
score += Math.max(0, Math.min(Math.floor(task.existingDiffLines / 50), 10));
}
if (task.changesPublicApi) score += 10;
if (task.changesDependencies) score += 10;
if (task.testScope === "integration") score += 5;
if (task.testScope === "full") score += 10;
score = Math.min(100, Math.max(0, score));
const taskType: TaskType =
score < 15
? "trivial_edit"
: score < 40
? "localized_fix"
: score < 70
? "multi_file_refactor"
: "architecture_synthesis";
const candidate = {
taskType,
complexityScore: score,
confidence: Math.max(0.2, 0.9 - missing.length * 0.35),
highRisk: Boolean(task.changesPublicApi || task.changesDependencies),
fallbackReason:
missing.length > 0 ? `missing metadata: ${missing.join(", ")}` : undefined,
};
const parsed = TaskComplexitySchema.safeParse(candidate);
return parsed.success ? parsed.data : CONSERVATIVE_FALLBACK;
}
const SUPPORTED_LEVELS = {
"gemini-3.8-flash": ("low", "medium", "high"),
} as const;
type SupportedModel = keyof typeof SUPPORTED_LEVELS;
type ThinkingLevel = (typeof SUPPORTED_LEVELS)(SupportedModel)(number);
const TIER_LEVELS: Record<TaskType, ThinkingLevel> = {
trivial_edit: "low",
localized_fix: "medium",
multi_file_refactor: "medium",
architecture_synthesis: "high",
};
const CONFIDENCE_FLOOR = 0.6;
export function resolveThinkingLevel(
model: SupportedModel,
c: TaskComplexity,
): ThinkingLevel {
let level: ThinkingLevel = TIER_LEVELS(c.taskType);
if (c.taskType === "localized_fix" && c.complexityScore < 30 && !c.highRisk) {
level = "low";
}
if (c.confidence < CONFIDENCE_FLOOR) level = "medium";
if (c.highRisk && level === "low") level = "medium";
const allowed: readonly string() = SUPPORTED_LEVELS(model);
if (!allowed.includes(level)) {
throw new Error(`Thinking level "${level}" is not supported by ${model}`);
}
return level;
}
interface RoutePlan {
classification: TaskComplexity;
level: ThinkingLevel;
}
interface UsageSnapshot {
inputTokens: number | null;
outputTokens: number | null;
thoughtTokens: number | null;
cachedTokens: number | null;
toolUseTokens: number | null;
}
interface AgentTaskResult {
plan: RoutePlan;
outputText: string | undefined;
usage: UsageSnapshot | null;
}
const REQUEST_TIMEOUT_MS = 120_000;
function toCount(v: unknown): number | null {
return typeof v === "number" && Number.isFinite(v) ? v : null;
}
function withTimeout<T>(p: Promise<T>, ms: number): Promise<T> {
let timer: ReturnType<typeof setTimeout> | undefined;
const timeout = new Promise<never>((_, reject) => {
timer = setTimeout(() => reject(new Error(`request timed out after ${ms} ms`)), ms);
});
return Promise.race((p, timeout)).finally(() => clearTimeout(timer));
}
class TaskRunError extends Error {
constructor(
message: string,
readonly plan: RoutePlan,
readonly original: unknown,
) {
super(message, { cause: original });
this.name = "TaskRunError";
}
}
export async function runAgentTask(
ai: GoogleGenAI,
model: SupportedModel,
task: CodingTask,
timeoutMs: number = REQUEST_TIMEOUT_MS,
): Promise<AgentTaskResult> {
const classification = classifyTask(task);
const level = resolveThinkingLevel(model, classification);
const plan: RoutePlan = { classification, level };
try {
const interaction = await withTimeout(
ai.interactions.create({
model,
input: task.instruction,
generation_config: { thinking_level: level },
store: false,
}),
timeoutMs,
);
const u = interaction.usage;
return {
plan,
outputText: interaction.output_text,
usage: u
? {
inputTokens: toCount(u.total_input_tokens),
outputTokens: toCount(u.total_output_tokens),
thoughtTokens: toCount(u.total_thought_tokens),
cachedTokens: toCount(u.total_cached_tokens),
toolUseTokens: toCount(u.total_tool_use_tokens),
}
: null,
};
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
throw new TaskRunError(message, plan, err);
}
}
interface TaskMetrics {
taskId: string;
model: SupportedModel;
taskType: TaskType | null;
thinkingLevel: ThinkingLevel | null;
classificationFallback: string | null;
status: "ok" | "empty_output" | "rate_limited" | "error";
latencyMs: number;
outputChars: number | null;
inputTokens: number | null;
outputTokens: number | null;
thoughtTokens: number | null;
cachedTokens: number | null;
toolUseTokens: number | null;
estimatedCostUsd: number | null;
pricingBasis: string;
errorMessage?: string;
}
const INPUT_USD_PER_M = 0.75;
const OUTPUT_USD_PER_M = 3.75;
const PRICING_BASIS = `gemini-3.8-flash paid standard tier, $${INPUT_USD_PER_M}/M input, $${OUTPUT_USD_PER_M}/M output incl. thinking, valid through 2026-12-31`;
const PRICING_VALID_UNTIL = Date.parse("2026-12-31T23:59:59Z");
const OUTPUT_INCLUDES_THOUGHTS = false;
export function estimateCostUsd(
u: UsageSnapshot | null,
now: number = Date.now(),
): number | null {
if (!u) return null;
const { inputTokens, outputTokens, thoughtTokens, cachedTokens, toolUseTokens } = u;
if (inputTokens === null || outputTokens === null || thoughtTokens === null) return null;
if (cachedTokens === null || toolUseTokens === null) return null;
if (cachedTokens > 0 || toolUseTokens > 0) return null;
if (now > PRICING_VALID_UNTIL) return null;
const billedOutput = OUTPUT_INCLUDES_THOUGHTS ? outputTokens : outputTokens + thoughtTokens;
return (
(inputTokens * INPUT_USD_PER_M) / 1_000_000 +
(billedOutput * OUTPUT_USD_PER_M) / 1_000_000
);
}
export function isRateLimited(err: unknown): boolean {
const status = (err as { status?: unknown } | null)?.status;
const message = err instanceof Error ? err.message : String(err);
return status === 429 || status === "429" || message.includes("RESOURCE_EXHAUSTED");
}
function logMetrics(m: TaskMetrics): void {
console.log(JSON.stringify({ ts: new Date().toISOString(), ...m }));
}
export async function trackedRun(
ai: GoogleGenAI,
model: SupportedModel,
task: CodingTask,
timeoutMs: number = REQUEST_TIMEOUT_MS,
): Promise<TaskMetrics> {
const start = performance.now();
try {
const result = await runAgentTask(ai, model, task, timeoutMs);
const metrics: TaskMetrics = {
taskId: task.id,
model,
taskType: result.plan.classification.taskType,
thinkingLevel: result.plan.level,
classificationFallback: result.plan.classification.fallbackReason ?? null,
status: result.outputText?.trim() ? "ok" : "empty_output",
latencyMs: performance.now() - start,
outputChars: result.outputText?.length ?? null,
inputTokens: result.usage?.inputTokens ?? null,
outputTokens: result.usage?.outputTokens ?? null,
thoughtTokens: result.usage?.thoughtTokens ?? null,
cachedTokens: result.usage?.cachedTokens ?? null,
toolUseTokens: result.usage?.toolUseTokens ?? null,
estimatedCostUsd: estimateCostUsd(result.usage),
pricingBasis: PRICING_BASIS,
};
logMetrics(metrics);
return metrics;
} catch (err) {
const plan = err instanceof TaskRunError ? err.plan : null;
const original = err instanceof TaskRunError ? err.original : err;
const metrics: TaskMetrics = {
taskId: task.id,
model,
taskType: plan?.classification.taskType ?? null,
thinkingLevel: plan?.level ?? null,
classificationFallback: plan?.classification.fallbackReason ?? null,
status: isRateLimited(original) ? "rate_limited" : "error",
latencyMs: performance.now() - start,
outputChars: null,
inputTokens: null,
outputTokens: null,
thoughtTokens: null,
cachedTokens: null,
toolUseTokens: null,
estimatedCostUsd: null,
pricingBasis: PRICING_BASIS,
errorMessage: (err instanceof Error ? err.message : String(err)).slice(0, 500),
};
logMetrics(metrics);
return metrics;
}
}
const MODEL: SupportedModel = "gemini-3.8-flash";
const TASKS: CodingTask() = (
{
id: "rename-greet",
operation: "rename",
affectedFiles: 1,
testScope: "unit",
changesPublicApi: false,
changesDependencies: false,
instruction:
"Rename the parameter `usr` to `user` in this function and return only the updated function:
function greet(usr) { return 'Hi ' + usr.name; }",
},
{
id: "extract-validation",
operation: "refactor",
affectedFiles: 4,
testScope: "unit",
changesPublicApi: false,
changesDependencies: false,
instruction:
"Describe a step-by-step plan to extract duplicated email validation logic from four TypeScript route handlers into a shared module, keeping existing behavior unchanged.",
},
{
id: "billing-plugins",
operation: "design",
affectedFiles: 14,
changesPublicApi: true,
changesDependencies: true,
testScope: "full",
instruction:
"Propose a plugin architecture for moving billing logic out of a monolithic Express application, including module boundaries, the public plugin interface, and a migration sequence that keeps existing endpoints working.",
},
);
const human = new Console(process.stderr);
async function main(): Promise<void> {
const apiKey = process.env.GEMINI_API_KEY;
if (!apiKey) {
console.error("GEMINI_API_KEY is not set. Export it before running.");
process.exitCode = 1;
return;
}
const ai = new GoogleGenAI({ apiKey });
const results: TaskMetrics() = ();
for (const task of TASKS) {
const m = await trackedRun(ai, MODEL, task);
results.push(m);
if (m.status === "rate_limited") {
console.error("Rate limited; skipping remaining tasks.");
break;
}
}
human.table(
results.map((m) => ({
task: m.taskId,
tier: m.taskType,
level: m.thinkingLevel,
status: m.status,
latencyMs: Math.round(m.latencyMs),
input: m.inputTokens,
thought: m.thoughtTokens,
output: m.outputTokens,
costUsd: m.estimatedCostUsd?.toFixed(6) ?? "n/a",
})),
);
if (results.length < TASKS.length || results.some((m) => m.status !== "ok")) {
process.exitCode = 1;
}
}
const isEntry =
process.argv(1) !== undefined && import.meta.url === pathToFileURL(process.argv(1)).href;
if (isEntry) {
main().catch((err: unknown) => {
console.error(err instanceof Error ? err.message : String(err));
process.exitCode = 1;
});
}You can test the routing and metrics logic without network access. Save the following as src/thinking-router.test.ts. It uses only the built-in node:test and node:assert modules, stubs the Interactions client, and runs with npm test.
import { test } from "node:test";
import assert from "node:assert/strict";
import type { GoogleGenAI } from "@google/genai";
import {
classifyTask, resolveThinkingLevel, estimateCostUsd, isRateLimited,
trackedRun, CONSERVATIVE_FALLBACK,
} from "./thinking-router.js";
const M = "gemini-3.8-flash" as const;
const full = { changesPublicApi: false, changesDependencies: false };
test("local rename with full metadata => trivial/low", () => {
const c = classifyTask({ id: "a", operation: "rename", instruction: "", affectedFiles: 1, testScope: "unit", ...full });
assert.equal(c.complexityScore, 5);
assert.equal(resolveThinkingLevel(M, c), "low");
});
test("rename with undeclared API/deps flags never routes low", () => {
const c = classifyTask({ id: "b", operation: "rename", instruction: "", affectedFiles: 1, testScope: "unit" });
assert.ok(c.confidence < 0.6);
assert.equal(resolveThinkingLevel(M, c), "medium");
});
test("NaN input => conservative fallback => medium; negative diff cannot lower score", () => {
const c = classifyTask({ id: "c", operation: "bug_fix", instruction: "", affectedFiles: Number.NaN, testScope: "unit", ...full });
assert.deepEqual(c, CONSERVATIVE_FALLBACK);
assert.equal(resolveThinkingLevel(M, c), "medium");
const d = classifyTask({ id: "d", operation: "bug_fix", instruction: "", affectedFiles: 1, testScope: "unit", existingDiffLines: -1000, ...full });
assert.equal(d.complexityScore, 25);
});
test("cost: known counters, missing counters, expiry", () => {
const u = { inputTokens: 1000, outputTokens: 200, thoughtTokens: 800, cachedTokens: 0, toolUseTokens: 0 };
const inWindow = Date.parse("2026-10-01T00:00:00Z");
assert.ok(Math.abs(estimateCostUsd(u, inWindow)! - 0.0045) < 1e-12);
assert.equal(estimateCostUsd({ ...u, thoughtTokens: null }, inWindow), null);
assert.equal(estimateCostUsd(u, Date.parse("2027-01-02T00:00:00Z")), null);
});
test("isRateLimited detects status and message", () => {
assert.equal(isRateLimited({ status: 429 }), true);
assert.equal(isRateLimited(new Error("RESOURCE_EXHAUSTED: quota")), true);
assert.equal(isRateLimited(new Error("boom")), false);
});
test("integration: trackedRun statuses with stubbed Interactions client", async () => {
const stub = (create: () => Promise<unknown>) => ({ interactions: { create } }) as unknown as GoogleGenAI;
const task = { id: "t", operation: "rename" as const, instruction: "x", affectedFiles: 1, testScope: "unit" as const, ...full };
const usage = { total_input_tokens: 40, total_output_tokens: 20, total_thought_tokens: 10, total_cached_tokens: 0, total_tool_use_tokens: 0 };
const ok = await trackedRun(stub(async () => ({ output_text: "fn", usage })), M, task);
assert.equal(ok.status, "ok");
assert.equal(ok.thinkingLevel, "low");
assert.equal(ok.thoughtTokens, 10);
const empty = await trackedRun(stub(async () => ({ output_text: "", usage })), M, task);
assert.equal(empty.status, "empty_output");
const zero = await trackedRun(stub(async () => ({ output_text: "fn", usage: {} })), M, task);
assert.equal(zero.estimatedCostUsd, null);
const rl = await trackedRun(stub(async () => { throw Object.assign(new Error("RESOURCE_EXHAUSTED"), { status: 429 }); }), M, task);
assert.equal(rl.status, "rate_limited");
assert.equal(rl.thinkingLevel, "low");
const hang = await trackedRun(stub(() => new Promise(() => {})), M, task, 50);
assert.equal(hang.status, "error");
assert.match(hang.errorMessage ?? "", /timed out/);
});As a sanity check after npm run build, the following command imports the built module without making any API calls and prints the level for a fully specified local rename. It should print low.
node --input-type=module -e "import('./dist/thinking-router.js').then(m=>console.log(m.resolveThinkingLevel('gemini-3.8-flash', m.classifyTask({id:'x',operation:'rename',instruction:'',affectedFiles:1,testScope:'unit',changesPublicApi:false,changesDependencies:false}))))"The table below separates assigned settings from observed outcomes. The classifier logic fixes the assigned columns. Only an executed run can fill the observed ones: token counts and latency vary between runs and do not follow from a tier label. Live output, usage, availability, and pricing all depend on external systems, so nobody can reproduce them exactly from this article.
Illustrative: observed columns require a live run
| Task | Tier (assigned) | Score | Level (assigned) | Latency | Input / Thought / Output tokens | Est. cost (pricing as of Sep 28, 2026) |
|---|---|---|---|---|---|---|
| rename-greet | trivial_edit | 5 | low | observed at runtime | observed at runtime | computed from usage |
| extract-validation | multi_file_refactor | 54 | medium | observed at runtime | observed at runtime | computed from usage |
| billing-plugins | architecture_synthesis | 100 | high | observed at runtime | observed at runtime | computed from usage |
Benchmark Methodology: Thinking Level vs. Accuracy vs. Cost
No benchmark results exist for this router. The outline supplied no sample sizes, run logs, correctness judgments, or completed-call costs, so this section offers a methodology and a blank template instead of numbers. The evidence here establishes neither „diminishing returns on simple tasks” nor „high thinking is necessary for architecture.” Do not assume architecture tasks need high because of product positioning. A lower-token call is also not necessarily cheaper per successful fix, because a failed patch can trigger extra repair calls.
SWE-bench offers a defensible pattern for judging correctness. Its evaluation applies a patch inside a repository environment and runs the tests, and its Verified subset contains 500 human-validated instances (per the official SWE-bench documentation). That makes it a model for evaluation methodology. It provides no measurements of this router and no proof that any one Gemini level is best.
A controlled comparison should:
- Hold the model ID, router code, tool permissions, prompts, test environment, and task set constant.
- Run five arms: the default,
low,medium,high, and the router. - Record actual
interaction.usage, wall-clock latency, end-to-end agent time, classifier overhead, failure and retry rates, test results, and cost per successfully completed task. - Stratify by task type and scope.
- Count a patch as correct only when it resolves the task and existing behavior still passes regression tests.
A reproducible report states the model version, run dates, pricing tier, and whether it included cached or tool-use tokens.
| Task Complexity Tier | Thinking Level | Tasks / Runs | Avg Latency (with distribution) | Avg Tokens | Cost per Call | Cost per Successful Task | Code Correctness Rate |
|---|---|---|---|---|---|---|---|
| Trivial edits | not measured | ||||||
| Localized fixes | not measured | ||||||
| Multi-file refactors | not measured | ||||||
| Architecture synthesis | not measured |
Calibration Best Practices for Production Coding Agents
Validating Supported Thinking Levels and Spend Limits
For gemini-3.8-flash, the „clamp” is the allowlist described earlier. Keep medium as the control arm when testing whether routing helps a real workload.
Monitor actual thought and output token usage, not just the selected level. Rate limits cover requests, input tokens, and spend. Keep 429 RESOURCE_EXHAUSTED separate from task or model failures in telemetry, because each needs a different response: a spend-limit failure calls for backoff, a failed patch for a different route or a retry. The sample stops on a spend-limit failure but implements no retry or backoff policy.
For legacy 2.5 models, a numeric budget clamp is a separate concern, and the requested value does not guarantee actual consumption.
Re-calibrating Based on Live Telemetry
Trigger recalibration when correctness, total task cost, tail latency, model ID, or pricing changes. Each team must set workload-specific thresholds for those signals from its own measured baseline. A rise in thought tokens alone is not enough. The pricing change scheduled for January 1, 2027 is one such trigger.
Keep a rollback path. If a lower-effort route starts producing more failed patches, revert that tier to the default.
When to Bypass Classification Entirely (Trusted Task Templates)
A trusted, narrowly specified task template can skip an LLM classifier and use a fixed level, but only after tests validate that level for the template. It should never skip patch validation, tests, or approval checks. Google advises applications to validate function calls before executing them.
Multi-turn agents that carry conversation history through previous_interaction_id need stored interactions. That conflicts with the example’s store: false setting, so decide the privacy and state trade-off deliberately before copying a single-call configuration into a multi-turn agent.
Treat Thinking Level as a Measured Runtime Choice
The loop is: classify the task, select a model-supported thinking level, call the Interactions API, measure request outcomes, and separately measure complete-task outcomes after patch application and verification. On the recommended gemini-3.8-flash path, the runtime variable is thinking_level, not a numeric thinkingBudget. Gemini’s medium default is the baseline the router has to beat.
Gemini’s
mediumdefault is the baseline the router has to beat.
Adapt thinking-router.ts to your agent’s real task distribution. Adjust operation scores, tier thresholds, and fallback rules, then validate each change against tests.
Adaptive or self-tuning level selection is a natural extension. It should adjust routes only on a verified success signal, meaning passing tests and accepted patches weighed against cost, and never on token savings alone.


