# Kelet
> Kelet continuously diagnoses why your LLM apps and AI agents fail in production, and generates the fix.
## What is Kelet?
Kelet is an AI agent that automatically investigates failures in your production AI agents/LLM Apps, runs a root cause analysis across many sessions, and generates targeted fixes — without manual trace debugging. It is not a monitoring or observability tool; it does the investigation and fixing for you.
Key capabilities:
- Automatic failure detection across thousands of production agent sessions
- Root cause analysis without manual trace scrolling
- Prompt patch generation with before/after reliability measurements
- Works with OpenTelemetry, Langfuse, Mixpanel, OpenAI, Anthropic, LangChain, CrewAI, and more
- SOC 2 certified, free to start
## Pages
- [Homepage](https://kelet.ai/): Product overview, how it works, and FAQ
- [Pricing](https://kelet.ai/pricing.md): Plans, features per tier, and pricing FAQ
- [Privacy](https://kelet.ai/privacy): Privacy policy
## How It Works
- **We collect your traces and signals.** Connect to your agent's stack in minutes. Every agent interaction, signal, and user feedback flows in automatically.
- **Know exactly why your agent failed.** Kelet reads every trace so you don't have to. Root causes surface in minutes, backed by evidence — not gut feeling.
- **Ship the fix. Know it held.** From root cause to prompt patch — with before/after reliability measurements. Kelet runs the patch against real sessions and shows you the before/after. No more flying blind after a deploy.
## How Kelet Compares
| Capability | Kelet | Langfuse | Arize | Datadog | Langsmith | Braintrust | Raindrop |
|-----------------------------------------|-------|----------|---------|---------|-----------|------------|----------|
| Trace collection | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Trace viewer for engineers | No | Yes | Yes | Yes | Yes | Yes | Yes |
| Agent that does the investigation | Yes | No | No | No | No | No | No |
| Automatic failure detection | Yes | No | Partial | No | No | No | No |
| Root cause analysis | Yes | No | No | No | No | No | No |
| Automated fix | Yes | No | No | No | No | No | No |
| Signal/score collection | Yes | Partial | Partial | No | Partial | Partial | Yes |
| Free tier | Yes | Yes | Yes | No | Yes | Yes | No |
## FAQ
**What does Kelet actually do?**
Kelet reads your production AI agent traces and signals, clusters failure patterns across thousands of sessions, and surfaces root causes with evidence — so you ship fixes instead of hypotheses. Think of it as a detective that investigates every failure automatically.
**What kinds of AI agents and LLM applications does Kelet work with?**
Any agent or LLM application where you own the code — agentic loops, multi-step workflows, RAG pipelines, chatbots, autonomous agents. If you built it and you ship it, Kelet can help you improve it. That includes agents built with LangChain, LangGraph, Google ADK, PydanticAI, Mastra, CrewAI, AutoGen, LlamaIndex, Haystack, Semantic Kernel, or directly on the OpenAI, Anthropic, Gemini, or LiteLLM APIs.
Two situations where Kelet is not the right fit: If you use AI tools built by others (Cursor, Claude Code, Copilot as a developer), you're a user, not a builder — Kelet isn't designed for your use case. Similarly, if you're building a skill or plugin inside an existing agentic platform, you're extending infrastructure you don't control, and Kelet can't instrument that.
But if you're building your own agent using any LLM SDK or framework — you own that agent, and Kelet is exactly for you.
**How long does integration take?**
Five minutes. Install via the Kelet installer skill — or `pip install kelet` / `npm install kelet` if you prefer to do it manually — add two lines to your agent code, and traces start flowing. Kelet is fully OpenTelemetry-compliant — any OTEL-instrumented agent works out of the box, no infrastructure changes needed.
**Where does Kelet actually run?**
On Kelet's servers. Once you install Kelet — via the SDK or the installer skill — traces and signals start flowing to our infrastructure automatically. It's SOC 2 certified and runs 24/7, continuously ingesting your traces, finding failure patterns, building hypotheses, and proposing targeted fixes.
The LLM tokens powering that analysis don't touch your model API bill — Kelet covers them. You pay Kelet based on usage. See kelet.ai/pricing.
**Is Kelet a skill or a service?**
A service. Kelet is an agent that runs on Kelet's servers around the clock — not a plugin you invoke, not something you run manually. The installer skill is just how you connect it. Once connected, Kelet works continuously: reading your traces, clustering failure patterns across thousands of sessions, building root cause hypotheses, and proposing targeted fixes. You don't run it. It runs for you.
**What are "signals" and why do they matter?**
Signals are probabilistic hints that something went wrong in a session: a thumbs-down rating, a user editing AI output, an abandoned conversation, or a synthetic LLM-as-judge check you configure. They tell Kelet where to look in your traces — not verdicts, but clues that guide the investigation.
**How is Kelet different from Langfuse, Arize, Logfire, or other observability tools?**
Those tools show you traces. Kelet reads them for you. Observability platforms are thermometers — they report symptoms. Kelet is the doctor that diagnoses root causes and generates targeted prompt patches. You no longer need to scroll thousands of traces manually.
**How does Kelet actually find root causes?**
Kelet works like a detective. Every session leaves a trail — LLM calls, tool invocations, retrieval steps, every agent hop. Kelet uses signals as clues: a thumbs-down, an edited AI response, an abandoned conversation, a synthetic LLM-judge flag. It follows each thread through your traces, cross-references patterns across thousands of sessions, and builds a root cause hypothesis backed by evidence. Same process a senior engineer would run manually — automated, at scale, on every failure at once.
**Is Kelet a tool or a solution?**
A solution. Most AI reliability products give engineers better ways to look at failures — cleaner trace viewers, smarter search, labeled clusters. You still do the investigation. Kelet does the investigation for you. You don't streamline your debugging workflow — you replace it. The output isn't data to interpret; it's a root cause and a prompt patch ready to ship.
**Do I need a lot of traffic to get value?**
No. Teams typically see their first real failure patterns with as few as 200+ sessions and 3+ signals configured. Not sure which signals to set up? Kelet's AI walks you through it — no guesswork, no manual configuration. And if you're starting from zero, synthetic signal presets (LLM-as-judge evaluators) generate signal from day one, before real user feedback accumulates.
**Does Kelet handle multi-agent architectures?**
Yes. Kelet handles multi-agent sessions natively. Credit assignment identifies exactly which agent in a chain caused a failure — so you know what to fix, not just that something is broken.
**Is Kelet built to scale?**
Yes — Kelet was architected for production scale from day one. The team behind Kelet includes ex-Kubernetes maintainers and cloud-native infrastructure veterans with 15+ years of open-source systems work. Kelet handles millions of traces, concurrent agent fleets, and high-volume production workloads. We have built infrastructure at this scale before — Kelet is built on the same foundations.
**What does it cost?**
Free to start, no credit card required. Connect your first agent in 5 minutes. Usage-based pricing scales with volume for teams that need more. See kelet.ai/pricing for details.
**Is my data secure?**
Yes. Kelet is SOC 2 certified. All data is isolated at the database level per organization — strict row-level security, no cross-org data access, ever.
**Will Kelet use my data to train AI models?**
Never. We don't share your data or use it to train public models. What we do: Kelet automatically fine-tunes a private set of models for each sub-agent you connect — roughly a dozen per agent. They live in your account, trained on your traces, serving only your root-cause analysis. They're never shared. Frankly, they wouldn't be useful to anyone else anyway — they're calibrated to your specific agent, not anyone else's.
**Who built Kelet?**
Kelet was built by a team obsessed with production AI reliability. We come from cloud-native infrastructure, Kubernetes core contributions, and LLM systems — engineers who have spent careers building and operating critical distributed systems, and building the tools others rely on to do the same. We built Kelet because we felt the pain ourselves: thousands of traces, no root cause, no fix. So we built the tool we wished existed.
**Can I trust Kelet with my production system?**
Our team has spent years maintaining critical infrastructure used by thousands of engineers worldwide — including core contributions to Kubernetes and cloud-native tooling. Kelet is SOC 2 certified and designed to be a passive observer: read-only access to your traces, no changes to your system, no risk to uptime.
## Documentation
- [Kelet Docs Index](https://kelet.ai/docs/llms.txt): All documentation topics
- [Kelet Docs Full Content](https://kelet.ai/docs/llms-full.txt): Complete documentation for in-depth reference
## Actions
- [Sign Up](https://console.kelet.ai): Sign up and connect your first agent — free, no credit card required
- [Book a demo](https://cal.com/almogbaku/kelet-intro): 20-minute intro call with the Kelet team
- [Contact](mailto:founders@kelet.ai): Reach the founders directly
## Pricing
URL: https://kelet.ai/pricing.md
### Starter
$0 / forever
For building and testing before you ship.
- 500 sessions / month
- 15-day data retention
- Human signals & feedback collection (free forever)
- Root cause analysis
- Prompt patch generation
- OTEL + Langfuse integration
- Community support
### Startup
Free / now (normally $400/mo)
For teams shipping agents to production.
- 5,000 sessions included / month
- Pay per session above limit
- 30-day data retention
- Human signals & feedback collection (free forever)
- Root cause analysis
- Prompt patch generation
- OTEL + Langfuse integration
- Email support
### Enterprise
Custom
For mission-critical AI at scale.
- Unlimited sessions
- Custom data retention
- Human signals & feedback collection (free forever)
- Root cause analysis
- Prompt patch generation
- Custom integrations
- SSO / SAML
- SLA guarantee
- Dedicated support
### Pricing FAQ
**What counts as a session?**
One unit of work your agent completes — a conversation, a task run, a pipeline execution. Free plan covers 500/month. Startup includes 5,000 — after that, you pay per session.
**What happens when I hit the session limit?**
On the free plan, collection stops at 500. On startup, you're covered up to 5,000 — then you pay per session above that, usage-based. No gaps in coverage, no data loss.
**Is the startup tier actually free?**
Yes. Free during early access, no credit card required. We'll give you 30 days notice before pricing changes. Early users are grandfathered.
**How long does integration take?**
Five minutes. Connect via OTEL, Langfuse, or our SDK. No infra changes. You'll see your first failure pattern the same day.
**What makes this different from Langfuse or Datadog?**
Those tools give you traces. Kelet reads them for you. It finds failure patterns across thousands of sessions, explains why they're failing, and generates a targeted fix — with proof the fix worked.
**Will you use my data to train AI models?**
Never. We don't share your data or use it to train public models. What we do: Kelet automatically fine-tunes a private set of models for each sub-agent you connect — roughly a dozen per agent. They live in your account, trained on your traces, serving only your root-cause analysis. They're never shared. Frankly, they wouldn't be useful to anyone else anyway — they're calibrated to your specific agent, not anyone else's.
## Blog
### We built a durable agent that debugs durable agents
URL: https://kelet.ai/blog/temporal-agentic-monitoring.md
Date: 2026-06-18
Author: Almog Baku
_Originally published on the [Temporal blog](https://temporal.io/blog/we-built-a-durable-agent-debugs-durable-agents)._
_When regular software breaks, a harness agent (Claude Code, Codex, whatever) follows the stacktrace and
fixes it. AI fails silently across hundreds or thousands of traces, with no stack trace to follow. That's why
we had to stop building harnesses and start building proper, durable workflows._
**TL;DR: Kelet is an AI that continuously diagnoses quality failures in AI agents, and it's built on Temporal.
This post covers the architecture, why it had to be structured this way, why a naive agent loop over your
traces cannot solve the problem, and how you can use Kelet to debug your own Temporal-based agents.**
---
Look at a single failing session in your trace viewer. You'll see what went wrong. You won't see _why it keeps
happening_.
AI failures don't show up like ordinary bugs. The same input succeeds ten times and fails on the eleventh, and
no two failures look quite alike. The actual root cause is a _fuzzy cluster_: a pattern that only becomes
visible across hundreds of sessions, when you start asking what the failures have in common.
We worked with an insurance team whose two-agent pipeline kept misclassifying claims. Every individual trace
looked like a scoring error in the second agent. The real problem was the first agent: it was stripping
chronological order from call transcripts, and in this particular insurance company, the _sequence of events_
determines the claim type. One session would have sent you patching the wrong agent entirely. It took hundreds
of sessions before the shape emerged.
That's the problem Kelet solves. And it's the reason we built it on Temporal.
---
## Durable orchestration, not an agent loop
Coding agents (like Claude Code, Codex, and even OpenClaw) can easily diagnose software bugs. Give it the trace
or the log, and it'll find and fix the exception. That part works.
Root causes for AI Quality are the harder case, because they don't live in an individual session or a trace.
They emerge from the overlap pattern across many occurrences. To find that overlap, you need to do three
things: process each session as it arrives, accumulate hypotheses about what's going wrong, and then reason
across the accumulated set. Those are three different jobs with their own latency profiles, inputs, and
coordination requirements, and you can't fold them into a single LLM call.
Dumping thousands of traces into one LLM call doesn't fix this either. The bottleneck isn't context length.
It's that pattern-matching across sessions requires building up state over time, gating the next stage on the
previous one, and surviving restarts in the middle. That's pipeline infrastructure, and that's what we needed
Temporal for. Specifically, we needed these Temporal primitives:
- **Long-running state:** Sessions accumulate over hours and days; analysis can't be synchronous.
- **Durable Execution:** Workers restart mid-analysis; each stage must resume cleanly from where it left off.
- **Event-driven coordination:** A new Signal should trigger reprocessing immediately, without polling
- **Cross-Workflow gating:** Cluster analysis only fires once enough hypotheses have accumulated.
A cron runner can't do the event-driven part. A generic task queue can't hold the durable long-running state.
We needed Workflow hierarchy, Signals, `wait_condition`, and `continue_as_new` (to reset Event History for
long-running workflows without losing state). Temporal gave us all four.
---
## How we built it
The algorithmic problem above drove the architecture directly. You need analysis stages that are isolated
(tractable input → tractable output), coordinated (each stage gates on the previous), and durable (any stage
can restart cleanly after a Worker crash).
Temporal gave us the primitives to build those properties without writing infrastructure from scratch. The
system that runs in production today is built around a four-level Workflow hierarchy.
```mermaid
graph TD
EV["Event Router Workflow (one per org/project · long-lived)"]
EV -->|new session signal| SW["Session Workflow (one per session · debounced 5 min)"]
subgraph SIG["Session Diagnosis · run in parallel"]
direction LR
CA[Signal Enrichment] ~~~ SE[Signal Merging] ~~~ IS[Agent Interrogation]
EM[Embeddings] ~~~ IN[Insights] ~~~ AG[Agent Assignment] ~~~ MORE[...]
end
SW --> SIG
SIG --> AA["Agent Aggregation Workflow (cross-session · per agent)"]
AA -->|per failure cluster| INV["Investigate Issue Workflow (root cause → prompt patch)"]
```
_This is a simplified view of what we actually run. The real system has more Workflow types, more nesting, and
additional coordination paths._
**Session Workflow** — one per session, but not triggered immediately. It waits for a 5-minute silence window
before starting analysis, avoiding partial-state work when a session is still in progress:
```python
DEBOUNCE_WINDOW: timedelta = timedelta(minutes=5)
# In the run loop:
await workflow.wait_condition(self._should_process, timeout=wait_time)
def _should_process(self) -> bool:
return (workflow.now() - self._last_message) >= self.DEBOUNCE_WINDOW
```
The Workflow rebuilds all state from the database on startup — no in-memory state that can't survive a Worker
restart.
**Signal Workflows** — Signal enrichment, merging, and agent interrogation run in parallel per session. They're
fully independent; there's no reason to serialize them.
**Agent Aggregation Workflow** — collects hypothesis attributions across sessions for a given agent. This is
where individual session analyses become a cross-deployment failure profile. The stage that solves the
second-order problem from earlier.
**Investigate Issue Workflow** — takes a failure cluster, reasons over it, produces a root cause with evidence,
and generates a prompt patch. This stage only fires once enough sessions have accumulated in the aggregation
layer, which is the gate that makes the reasoning tractable.
---
## We use Kelet on Kelet, and you can too
We monitor Kelet's own Temporal Workflows in production with Kelet. Temporal-native integration was the
starting point, not an afterthought — we needed a system that could work on itself.
The integration is a single Temporal plugin that does two things:
1. Configure Temporal's built-in `OpenTelemetryPlugin` so your traces are wired up out of the box.
2. Propagate Kelet's session attributes (session ID, user ID, metadata) across Workers, Workflows, and
Activities — including child Workflows and `continue_as_new` — by stamping Temporal headers on the way out
and reading them on the way in to attach as OTel span attributes.
```mermaid
flowchart LR
subgraph Out["Outbound interceptor"]
O[stamp session attrs
onto Temporal header]
end
subgraph In["Inbound interceptor (Activity)"]
I[read header →
open agentic_session]
end
Caller[Workflow / Client] --> Out
Out -->|header| T[(Temporal)]
T --> In
In --> Body[Activity body]
Body -. emits .-> Span[OTel span
session=… attr attached]
```
Headers become part of event history, so propagation replays deterministically and survives Worker restarts and
`continue_as_new` for free. On the Activity side, the inbound interceptor opens `agentic_session(...)` around the
body, and every OTel span emitted inside picks up the session as an attribute — your LLM calls, retrieval, and
tool spans land in the right Kelet session without your code ever touching it.
One detail that matters in our own prod: the interceptor filters out Kelet's own monitoring Workflows so they
don't get re-ingested as sessions. Otherwise self-monitoring becomes an infinite loop — every diagnosis spawns a
session, which spawns another diagnosis, which spawns another session.
If you're building agents on Temporal, the fastest path:
```bash
npx skills add Kelet-ai/skills
```
The skill registers the plugin for you. Hand-wired, it's three lines:
```python
from kelet.temporal import KeletPlugin
kelet.configure(api_key="...", project="my-agent")
client = await Client.connect("localhost:7233", plugins=[KeletPlugin()])
worker = Worker(client, task_queue="ai", workflows=[MyWorkflow], activities=[my_activity])
```
Point Kelet at your own Workflows and the durable analysis pipeline — cross-session aggregation, root-cause
clustering, prompt patches — runs on top of the traces you're already producing.
---
## What this looks like in production
We diagnose thousands of sessions per day. No human in the loop. No polling, no cron jobs, no manual triggers —
new sessions flow in via Temporal Signals, and the pipeline picks them up from there.
Worker restarts don't drop anything. Every Workflow resumes from where it stopped. The infrastructure we
_didn't_ have to build: a scheduler, a distributed state machine, per-stage retry logic, and a coordination
layer. Temporal handles all of that.
That durability rests on a contract with Temporal: activities are idempotent, and signals carry an idempotency
key so a worker restart that re-delivers a signal doesn't double-process the session.
The work that took us months was the analysis pipeline itself: getting the stage boundaries right, tuning the
debounce windows, figuring out how to aggregate across sessions, and getting the second-order reasoning to turn
thousands of individual diagnoses into a single, named root cause.
Temporal freed us to spend that time on the actual problem instead of on coordination plumbing. Kelet is what we
built with that time. If you're running agents or LLM applications on Temporal, pointing Kelet at them gets you
continuous Failure Analysis on top of the Durable Execution you already have, which is what surfaces the
cross-session patterns your traces won't show.
---
### We Spent 30% of Our Engineering Debugging AI Agents. So I Built One to Do It.
URL: https://kelet.ai/blog/launching-kelet.md
Date: 2026-04-12
Author: Almog Baku
**TL;DR:** Kelet is an AI agent that automates root cause analysis for production AI agents. It reads your traces across
thousands of sessions, identifies failure patterns no human would find manually, and generates validated prompt fixes
with before/after proof. [Free during beta](https://console.kelet.ai).
---
I've built 50+ LLM apps and agents over the past few years, some scaled to millions of transactions a day. And after
all of that, the thing I'm most frustrated by is still how we debug them. We have models writing production code,
passing exams, replacing entire workflows, and when they break in production? Our best strategy remains:
open the monitoring system, scroll, squint, guess, patch, deploy, _pray._ Monday morning, same thing again.
I've talked to over 112 AI engineers. Same story, every time. About a third of their week just... _gone._
Not building anything. Not shipping. Just _scrolling traces._ And it's not just anecdotal:
[McKinsey](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)
and [Gartner](https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025)
report that roughly 90% of AI projects look great in PoCs and collapse in production, because nobody has a reliable
way to diagnose what's going wrong at scale.
I tried everything. Evals that passed in staging, then fell apart the moment real users showed up. Autonomous
monitoring that caught basically nothing useful. Every observability dashboard I could find, all of which gave
me beautiful charts and _zero actual answers._
The only thing that actually worked was the oldest, least glamorous trick in data science: **error analysis.**
Going through production data, session by session, tagging failures by hand, slowly building a picture of what's
_actually_ breaking. The best AI engineers I know, people shipping real agents at serious scale, were all doing
the same thing. _Spreadsheets._ Every morning. One session at a time. It felt like 2015 all over again, manually
labeling training data.
It works. It's also an absurd waste of very expensive engineering time. And it doesn't scale.
## "Just throw Claude Code at it"
Everyone tried this. Claude Code, Cursor, deep research agents, autonomous debugging loops. Same wall every time.
Nobody in this industry seems willing to say it out loud, so I will: a single LLM cannot solve this problem.
It's the wrong shape of tool for the job.
In traditional software, a coding agent can follow a traceback straight to the root cause. One session, one bug,
one fix. AI failures don't work that way. The same input succeeds ten times, fails on the eleventh, works again
for no apparent reason. One hallucination is a data point, not a diagnosis. You need to observe a failure repeat
across _hundreds_ of sessions before you can separate real patterns from noise.
Think of it like dozens of needles scattered across thousands of haystacks — except the needles are _connected_ in
ways you can only see when you zoom out far enough. No single LLM call is going to surface that. You need an
assembly of specialized models (some LLM-based, some classical ML) learning continuously over weeks and months.
We train dozens of models _per subagent_ in your pipeline.
The question I kept coming back to: we're building agents for legal research, code generation, customer support.
_Why can't an agent do the error analysis itself?_ Let the humans make the judgment calls. Let the machine do the heavy lifting. So that's what we built.
## The failure nobody saw coming
An insurance company we worked with had a two-agent pipeline: the first agent summarizes support call transcripts,
the second classifies the claim type. Passed all their evals. Everyone was happy.
One session that stuck with me: a woman calls about a storm. Flooding, broken pipe, mud everywhere, the whole
disaster. The summarization agent pulls out the facts. The classification agent reads that summary and outputs
"water system problem." The human underwriter glances at it, nods, and quietly changes it to "weather event."
Doesn't open a ticket. Doesn't flag anything. Just silently corrects it and moves on.
If you looked at that one session in your trace viewer, you'd blame the classification agent. Add a better
few-shot example to the prompt. Ship the fix. Done, right?
We analyzed _hundreds_ of their sessions: fire claims, power outages, electrical failures. The classification
agent wasn't the problem. The _summarization_ agent was. It kept flattening the timeline into a bag of facts
with no chronological order. In insurance, _the sequence of events_ is what determines the claim type, not the
events themselves. That's institutional knowledge that experienced underwriters carry in their heads. It's not
written down anywhere. A single session would have sent you debugging the wrong agent entirely.
> **Tip — The core insight**
> One session shows you that something went wrong. Hundreds show you _why_ it
> keeps going wrong. And once you find the actual root cause, the fix is usually much simpler.
## Your observability tool is a very expensive screenshot
I'll say what the Langfuse and LangSmith teams won't: trace collection is a solved problem. OpenTelemetry
commoditized it. Langfuse, LangSmith, Datadog, Braintrust, Arize. They're all showing you the same underlying
data with different interfaces on top. For example, Langfuse will show you latency, token count, cost — everything
except _why_ the agent failed. At this point the real competition is about who renders traces more beautifully.
And not one of them can tell you _why your agent broke._
The way I see the current stack:
| Layer | What it does | Status |
| ----------------------- | -------------------------------------------------- | -------------------------- |
| **Traces + metrics** | Collecting and viewing what happened | Solved. Commoditized. |
| **Root cause analysis** | Understanding _why_ it failed, with evidence | Every tool stops here. |
| **Automated fix** | Generating a validated fix with before/after proof | Nobody does this. |
The whole observability industry built thermometers — some of them very good thermometers, I'll give them that.
But a thermometer doesn't diagnose strep throat. It confirms what you already knew: something is wrong. Your
agent is failing. Your users are complaining. You didn't need a fancier dashboard to tell you that.
What's missing from the stack is what comes after the thermometer: an actual diagnosis, and a prescription.
That's the gap Kelet fills.
If you're already on Langfuse, keep it. Kelet pulls your traces directly from there. No re-instrumentation needed. We sit on top of what you already have.
## What Kelet actually does
Your traces are the raw material. User signals — thumbs-down clicks, retries, silent edits where a user quietly
fixes what your agent got wrong (e.g., rewrites a hallucinated sentence) — point Kelet toward where the real
problems are concentrated.
1. **Collect**: Traces and signals flow in via OpenTelemetry, Langfuse, or the Kelet SDK. Five minutes to connect.
2. **Investigate**: Kelet reads every session, clusters failures, and pinpoints which agent in your pipeline caused the problem.
3. **Fix**: For every root cause, Kelet generates a prompt patch, with before/after quality metrics to prove it works.
When a failure pattern keeps repeating across hundreds of sessions, Kelet surfaces it as a named finding, a
specific root cause backed by evidence, not a hunch. Not "I spent an hour scrolling and I _think_ it might be
the retrieval step."
For extra confidence before shipping a fix, there's [GEPA](https://arxiv.org/abs/2507.19457) optimization, essentially an evolutionary search
tested against your real production sessions. It shows you the improvement _before_ you deploy anything, so
you're not crossing your fingers after every prompt change.
Numbers so far:
I still find that 14.3-minute figure surprising, honestly. It used to take us weeks. (The teams where it's
slower are the ones where the failure hasn't had enough repetitions yet. 14 minutes assumes sufficient volume
for a pattern to emerge. Connect earlier, and the time improves.)
## Try it
Your agent is failing somewhere right now. Scrolling through traces one by one isn't going to find it.
If you're building agents and want to talk through any of this: [founders@kelet.ai](mailto:founders@kelet.ai).
_— Almog_
---