I was mid‑night on a Saturday when my on‑call pager blared. The AI‑powered support bot we’d just shipped was spitting out “I’m sorry, I don’t understand” for every user query, and our cloud‑bill dashboard showed a sudden $12 k jump in LLM token spend. Turns out a mis‑tagged retry loop kept firing OpenAI calls even when the primary request failed—so every retry added cost without being reflected in any of our dashboards. Fixing the bug stopped the bleed, but the incident taught me a hard lesson: you can’t optimise AI spend without observability built into every agent.
- Instrument each AI call with OpenTelemetry spans that carry token‑count and model metadata.
- Tag deployments using Harness “AI Service Blueprints” for per‑deployment cost attribution.
- Create dashboards that correlate cost, latency, and business outcomes.
- Sample high‑frequency agents to keep overhead < 2 % while preserving accuracy.
- Set context‑aware alerts that surface cost anomalies before they hit the budget.
Before you start: You need Harness NextGen 2026 (SRM + CCM modules), OpenTelemetry SDK 1.24 for Python, @opentelemetry/api 1.13 for JavaScript/Node.js, a Kubernetes 1.31 cluster with the Harness GitOps Agent installed, and basic knowledge of LangChain or LlamaIndex.
How to Monitor AI Agent Costs with Harness Telemetry (2026)
Monitor AI agent costs per deployment in Harness by instrumenting your agent code with OpenTelemetry spans tagged with cost metadata (e.g., tokens, model, provider). Harness Observability collects this telemetry, allowing you to create dashboards and set alerts for specific deployments. Integrate with Cloud Cost Management for precise FinOps reporting and anomaly detection.
Why Granular AI Agent Cost Monitoring Matters in 2026
The Shift from Pilots to Production Deployments
Two years ago most teams were still running AI “proof‑of‑concepts” inside notebooks. Today every checkout, every help‑desk ticket, and even some internal tooling runs an LLM‑backed agent in production. That scale turns a few dollars of token spend into a multi‑million‑dollar line item if you’re not watching it tightly.
The High Stakes of Uncontrolled AI Spending
Untracked token usage is the new “shadow‑IT”. A single mis‑configured loop can double the cost of a deployment overnight. According to a 2025 GigaOm case study, a fintech firm cut “unknown” LLM spend by 35 % in Q1 after introducing fine‑grained cost monitoring. Without that visibility, you’re flying blind while the CFO’s eyes widen.
How Harness Observability Gathers AI Cost Telemetry Data
Understanding Key Telemetry Signals (Spans, Metrics, Logs)
| Signal | What it tells you | Example in AI agents |
|---|---|---|
| **Span** | End‑to‑end request lifecycle, includes attributes | `llm.inference` span with `tokens_used=342`, `model=claude‑3‑sonnet` |
| **Metric** | Aggregated numeric data, useful for trends | `tokens_per_minute` gauge |
| **Log** | Unstructured events, good for error context | `LLM call failed: rate‑limit` |
Harness SRM ingests OTel spans, while the CCM module pulls raw cloud‑billing data from Azure OpenAI, Anthropic, etc. When you push cost attributes on spans, the platform joins them to the right billing line, giving you per‑deployment cost views.
Agent‑Level vs. Deployment‑Level Instrumentation
*Agent‑level* spans capture each LLM call inside a single LangChain or LlamaIndex agent. *Deployment‑level* spans wrap the whole request that may traverse multiple agents, queues, or microservices.
Both are needed: the former answers “how many tokens did this call use?” while the latter answers “which deployment incurred that spend?”
**Tip:** Use the `span.set_attribute` method for cost fields and propagate the context across async calls—this prevents double‑counting when agents call each other.
Step‑By‑Step Strategy for Tracking Costs Per Deployment
Before diving into the steps, make sure you’ve read our **[Setting up OpenTelemetry for Harness](https://nileshblog.tech/harness-gitops-agent/)** tutorial – it covers the collector config you’ll need.
Step 1: Define and Tag AI‑Specific Service Blueprints
In Harness, create a **Service Blueprint** called `ai‑support‑bot`. Add custom tags like `model=claude‑3‑sonnet` and `team=customer‑support`. These tags flow into every span generated from that service.
# Harness blueprint (YAML, Harness CLI v2.6)
service:
name: ai-support-bot
tags:
model: claude-3-sonnet
team: customer-support
Step 2: Enrich Telemetry with Cost Metadata
Add the following snippet to every LLM request wrapper. It records token count, model name, provider, and request latency.
# Python 3.12, opentelemetry-sdk 1.24
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor, OTLPSpanExporter
import time
trace.set_tracer_provider(TracerProvider())
otlp_exporter = OTLPSpanExporter(endpoint="http://harness-collector:4317")
trace.get_tracer_provider().add_span_processor(BatchSpanProcessor(otlp_exporter))
tracer = trace.get_tracer("ai.agent")
def invoke_llm(prompt: str) -> str:
start = time.time()
# pretend we call OpenAI here and get token usage
response, tokens = fake_openai_call(prompt)
duration_ms = (time.time() - start) * 1000
with tracer.start_as_current_span("llm.inference") as span:
span.set_attribute("model", "gpt-4o")
span.set_attribute("provider", "azure-openai")
span.set_attribute("tokens_used", tokens)
span.set_attribute("duration_ms", duration_ms)
span.set_attribute("deployment", "ai-support-bot")
return response
// Node.js 20, @opentelemetry/api 1.13
const { trace, context, propagation } = require("@opentelemetry/api");
const { OTLPTraceExporter } = require("@opentelemetry/exporter-trace-otlp-http");
const { BatchSpanProcessor, NodeTracerProvider } = require("@opentelemetry/sdk-trace-node");
const provider = new NodeTracerProvider();
const exporter = new OTLPTraceExporter({url: "http://harness-collector:4318/v1/traces"});
provider.addSpanProcessor(new BatchSpanProcessor(exporter));
provider.register();
const tracer = trace.getTracer("ai.agent");
async function invokeLLM(prompt) {
const start = Date.now();
const { response, tokens } = await fakeAnthropicCall(prompt);
const durationMs = Date.now() - start;
const span = tracer.startSpan("llm.inference", {
attributes: {
model: "claude-3-sonnet",
provider: "anthropic",
tokens_used: tokens,
deployment: "ai-support-bot",
},
});
span.end();
return response;
}
Step 3: Create AI Agent Performance Dashboards
In Harness SRM, add a **New Dashboard** with the following widgets:
- **Cost per Deployment** – a stacked bar showing `$` by `deployment` tag.
- **Tokens vs. Latency** – a scatter chart (`tokens_used` on X, `duration_ms` on Y).
- **Anomaly Heatmap** – uses the built‑in AIOps model to surface spikes.
The dashboards can be shared with finance and product teams, turning raw telemetry into actionable insight.
Step 4: Set Up Granular, Context‑Aware Alerts
Harness AIOps lets you write alert policies in a YAML DSL. Example: fire when token spend per minute exceeds the 95th percentile for the last 7 days.
# alert-policy.yaml (Harness AIOps 2026)
policy:
name: high-token-spend
condition:
type: threshold
metric: tokens_per_minute
operator: gt
value: ${{ quantile(95, "tokens_per_minute", window="7d") }}
actions:
- type: pagerduty
target: "#ai-cost-alerts"
tags:
- deployment: ai-support-bot
Code and Architecture Best Practices for Accurate Costs
**My take:** Most teams treat cost as an after‑thought. In production you must bake cost tags into every retry and fallback path; otherwise you’ll see “missing data” warnings in CCM for exactly the calls that hurt your budget the most.
Instrumenting for Python and JavaScript AI Agents
Both snippets above show the minimal span. A production‑ready implementation adds:
- **Retry Loop Instrumentation** – each retry gets its own span, flagged `retry=true`.
- **Fallback Span** – when a fallback LLM is called, add `fallback=true`.
# Adding retry awareness (Python)
from tenacity import retry, stop_after_attempt, wait_exponential
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=1, max=4))
def invoke_llm_with_retry(prompt):
with tracer.start_as_current_span("llm.inference.retry") as span:
span.set_attribute("retry", True)
return invoke_llm(prompt)
// Adding fallback awareness (Node.js)
async function invokeWithFallback(prompt) {
const primary = await invokeLLM(prompt);
if (!primary) {
const fallbackSpan = tracer.startSpan("llm.inference.fallback", {
attributes: { fallback: true, model: "gpt-3.5-turbo" },
});
const result = await invokeLLMFallback(prompt); // another model
fallbackSpan.end();
return result;
}
return primary;
}
Implementing Resilient Error and Fallback Instrumentation
When an LLM call fails due to rate limits, you should capture the error code **and** the cost (often zero tokens but still a billed request). Adding `error.type` and `error.message` attributes makes post‑mortem analysis trivial.
with tracer.start_as_current_span("llm.inference") as span:
try:
response, tokens = real_llm_call(prompt)
span.set_attribute("tokens_used", tokens)
except Exception as exc:
span.set_status(trace.Status(trace.StatusCode.ERROR, str(exc)))
span.set_attribute("error.type", type(exc).__name__)
span.set_attribute("tokens_used", 0) # billed request may still count
raise
Production Case Studies and 2026 Benchmark Insights
The **[Multi‑Agent FinOps: 5 Patterns to Cut AI Costs (2026)](https://nileshblog.tech/?p=6734)** post walks through patterns you’ll see replicated here.
Case Study: Reducing LLM Token Spend by 35 %
A large fintech rolled out a LangChain‑based fraud‑detect bot. By adding per‑span token tags and the alert policy above, they discovered a runaway `retry=true` loop that was re‑sending the same 2 k‑token prompt on every timeout. After fixing the loop and capping retries to 2, token spend dropped from 12 M to 7.8 M per month – a 35 % reduction.
Published 2026 AI Agent Cost Benchmarks
| Provider | Avg. token cost (USD) | Avg. latency (ms) | 95th‑pct variance |
|---|---|---|---|
| Azure OpenAI (gpt‑4o) | 0.00012 | 210 | ±12 % |
| Anthropic (claude‑3‑sonnet) | 0.00010 | 180 | ±9 % |
| Cohere (command‑r) | 0.00009 | 240 | ±15 % |
These numbers come from Harness AIOps’ aggregated dataset of > 3 k deployments in 2026. Use them as baselines when you size your budgets.
System Design Trade‑offs: Balancing Detail vs. Overhead
The Telemetry Sampling Decision
Sampling at 100 % gives perfect accuracy but can add ~2 % CPU overhead on a busy inference service. A pragmatic rule of thumb: **sample 1 in 10** for agents that run > 5 k calls per minute, and keep 100 % for low‑volume, high‑value agents (e.g., legal‑review bots). Harness collector respects the `sampling.priority` attribute, so you can decide per‑span.
Managing Latency from Cost‑Aware Instrumentation
Spans are exported asynchronously, but the context propagation path adds a handful of nanoseconds. If you see > 5 ms added per request, look for **blocking exporters** or **synchronous log writes**. Switching the exporter to **OTLP over gRPC** and enabling batch processing (default in the snippets) usually brings the overhead back under 2 %.
Outcome‑Driven Reporting and Cost Attribution
Correlating Cost with Business KPIs
Link the `deployment` tag to a business metric table (e.g., `tickets_resolved`). In Harness SRM you can create a **Composite Widget** that shows “Cost per ticket resolved” over time. The metric often reveals that a 5 % increase in token spend translates to a 0.8 % dip in CSAT – a concrete ROI argument for finops investment.
Effective Reporting for DevOps and Finance Teams
Export the `cost_per_deployment` metric to a CSV or pull it via the Harness GraphQL API. Finance can then mash it into their budgeting tool, while DevOps sees live alerts. A shared **Cost‑Health Report** (PDF generated nightly) keeps both sides aligned.
Common Errors & Fixes
**Error:** *“Missing cost attributes on spans”* *Symptom:* CCM shows “unknown LLM spend” for a deployment. *Why it happens:* The collector