calibration infrastructure for AI

The AI that knows
when it doesn't.

Every model can answer. Almost none can tell you when to trust the answer. idk scores every AI output with a calibrated confidence percentage — and returns an honest "idk" instead of a confident hallucination.

model-agnostic · any LLM integration · 2 lines abstentions · free audit ledger · built in
0
answers scored*
0
abstentions issued*
0
calibration accuracy*
0
median scoring latency*
"I don't know" is a featureoverconfidence is a bugcalibration > charismatrust ships with a numberhallucinations don't announce themselves — we doshrug responsibly ¯\_(ツ)_/¯every answer, scoredabstain > fabricateτ = you decide
01 · the problem

AI answers everything with the same swagger. Even when it's making things up.

Language models are trained to always produce an answer — never to weigh whether they should. The wrong answer and the right one arrive in the same confident voice. That's fine for trivia. It's catastrophic for medicine, law, finance, and code.

0

Times a raw LLM says "I'm not sure" unprompted, per thousand answers it should. Confidence is uniform; correctness isn't.

?

Your current visibility into which AI outputs in production are wrong. If you can't measure it, you're shipping it.

idk

The three-letter answer that would have prevented the fabricated citation, the phantom API, and the invented refund policy.

What's our refund policy for enterprise contracts signed before 2024?
RAW MODEL · SOUNDS SURE, ISN'TEnterprise contracts signed before 2024 are eligible for a full refund within 90 days, per section 4.2 of the standard agreement.
WITH IDK · CONFIDENCE 22% → ABSTAINEDidk — I can't verify that policy from grounded sources. Routing to a human. (Section 4.2 doesn't exist. Neither does the 90-day window.)
One of these answers becomes a lawsuit. The other becomes a support ticket.
02 · the product

One engine, three ways to ship honesty.

idk sits between any model and your users. Every output comes back with a calibrated confidence score, supporting evidence, and — below your threshold — a clean abstention.

{ }

idk engine

the api

A drop-in calibration API. Send any prompt + any model's answer; get back {answer, confidence, evidence, abstain}. Model-agnostic. Two lines of code.

[!]

idk verify

the firewall

Middleware for production AI. Audits every output in-flight, flags overconfident claims, blocks or reroutes answers below threshold, and keeps a confidence ledger for compliance.

?_

idk ask

the showcase

A consumer answer app where every response shows its confidence — and sometimes the answer is just "idk." Radical honesty as a habit-forming product.

03 · the method

Confidence isn't a vibe. It's a measurement.

The engine estimates uncertainty from independent signals, then calibrates the fused score against outcomes — so "85%" means right about 85% of the time.

01 / SAMPLE

Self-consistency

Resample the model and measure semantic agreement. Answers that wobble under resampling are answers you shouldn't trust.

02 / GROUND

Evidence check

Claims are decomposed and matched against retrieved sources. Unsupported claims drag the score down — hard.

03 / CALIBRATE

Score against reality

Raw signals are fused and mapped to real-world hit rates using domain-tuned calibration curves.

04 / DECIDE

Answer or abstain

Above threshold τ: answer + score + evidence. Below: a structured "idk" with a fallback route — human, retrieval, or retry.

reliability diagramfig. 01
perfect 0% 100% stated confidence → actual accuracy →
raw model (says 90%, right 55%) with idk (says 90%, right ~90%)
04 · live demo

The terminal. Wired to a real model. Scored in real time.

This is not a mockup. Connect an Anthropic API key and the engine runs a live two-pass check — draft, then calibrate — and returns a confidence percentage. Below threshold τ, it abstains. Ask it about the future and watch it shrug.

idk engine — live session · api.anthropic.com
offline τ 0.75
[boot]idk engine v1.0 — calibration terminal
[boot]two-pass mode: SAMPLE → CALIBRATE · abstention threshold τ=0.75
[auth]paste an Anthropic API key to go live. it stays in this tab's memory — never stored, never sent anywhere but api.anthropic.com.
idk>
demo runs client-side against the Anthropic API with your key · commands: /help /model /threshold /clear · confidence here is model-self-estimated — the production engine adds evidence grounding + resampling
05 · why it pays

What a company buys when it buys "idk."

Calibration isn't a nice-to-have on top of AI — it's the control layer that makes AI deployable where money and liability live.

B/01

A liability shield with receipts

Every output ships with a score, evidence, and a ledger entry. When a regulator, court, or customer asks "why did your AI say that?" — you answer with an audit trail, not a shrug. (We handle the shrugs.)

B/02

Automation you can actually turn up

Most teams cap AI autonomy because they can't tell good answers from bad. With per-route thresholds, high-confidence answers auto-send while risky ones route to humans — so automation rates climb without incident rates following.

B/03

Compliance, pre-packaged

Emerging AI rules (EU AI Act and successors) demand risk management, transparency, and human oversight for high-risk systems. An abstention log — what your AI declined to answer, and why — is the cleanest oversight artifact yet invented.

B/04

Trust that compounds into usage

Users who see honest confidence scores — including the occasional "idk" — trust the confident answers more, and come back. Admitted unknowns are marketing for every answer you do stand behind.

Exposure calculator

// what unscored AI answers cost you — illustrative model, tune the sliders to your reality
BAD ANSWERS SHIPPED / MO3,000
MONTHLY EXPOSURE$120,000
CAUGHT & REROUTED BY IDK$102,000
ANNUAL EXPOSURE AVOIDED$1,224,000
illustrative arithmetic, not a quote. your error rate is the scary slider.
06 · who runs idk

Anyone whose AI talks to customers, courts, or capital.

If a wrong answer costs more than an API call, you're the customer.

support & CX platforms

Support automation leads

"Our bot promised a refund policy that doesn't exist."

with idk → only ≥90% answers auto-send; the rest route to agents with the draft attached. Deflection up, incidents down.

legal & professional services

LegalTech & law firm innovation teams

"The brief cited a case that was never decided."

with idk → every citation is decomposed and verified against sources; unverifiable claims abstain before they reach a filing.

healthcare & life sciences

Clinical informatics & digital health

"The summary invented a dosage."

with idk → τ=0.99 on anything clinical; below it, structured escalation to a human — with the uncertainty reasons attached.

financial services

Risk & compliance officers

"We can't deploy AI we can't audit."

with idk → the confidence ledger turns every output into an auditable event: score, evidence, decision, route. Exam-ready.

AI product teams

Builders shipping LLM features

"One viral hallucination undid a quarter of trust-building."

with idk → two lines of code, per-route thresholds, and a UI primitive users learn to love: the confidence chip.

agents & automation

Teams deploying autonomous agents

"It didn't ask. It just booked the wrong thing."

with idk → pre-action confidence gates: agents check "how sure am I?" before they click buy. Below τ, they stop and say idk.

07 · pricing

Pay for answers. Abstentions are free.

We only charge when the engine is confident enough to let an answer through. When it says idk, so does your invoice.

Curious

$0/mo
  • 10k checks / month
  • idk engine API
  • Community calibration set
  • 2 domains
Start free
most popular · we're 96% sure

Confident

$499/mo
  • 1M checks / month
  • idk verify middleware
  • Custom thresholds per route
  • Confidence ledger + audit export
  • Domain-tuned calibration
Start trial

Certain*

Custom
  • Unlimited checks
  • On-prem / VPC deploy
  • Compliance reporting (EU AI Act-ready)
  • Dedicated calibration engineer
Talk to us

*Nothing is certain. We're the company that tells you that.

// the idk manifesto

We taught machines the three hardest words.

Overconfidence is a bug.File it. Fix it. Ship the patch.

"I don't know" is a feature.The most underrated string in computing.

Trust ships with a number.No score, no deploy.

Calibration beats charisma.85% should mean 85%.

Abstain, then escalate.A shrug with a plan attached.

Honesty compounds.Every admitted unknown buys the next answer credibility.

// go deeper

Uncertainty is infrastructure.

Our white paper lays out the full argument: why models bluff, how calibration works, and what an honest AI stack looks like.

Download the white paper (PDF) hello@idk.io