Editorial Subdomain
Public Reader Platform
Cost Discipline & Inference Telemetry · §9.2

Cost Architecture & Inference Ladder

Strict separation of three cost pools (Build, Run, Inference) with target blended AI cost under €0.02 per published event.

The Three Cost Pools

Pool 1: Build Spend
Coding Agent

Budget ceiling: €100.00 / mo. Completely isolated key; never touches production database.

Hard ceiling verified
Pool 2: Run Spend
Infrastructure

Postgres + pgvector single system of record. Zero Redis/Kafka in MVP. Tier 0 host: €15–30/mo.

Zero premature microservices
Pool 3: Inference Spend
0.0035

Measured: 0.7000 / 1,000 events (Cap: €20.00).

Under €0.02 / event target
Five-Rung Inference Ladder (ADR-0003)Escalation only upon deterministic check failure
L0
Deterministic

Regex, entity dictionaries, heuristics

Cost: €0.0045 calls (52%)
L1
Local CPU

Local tokenizers & embeddings on CPU

Cost: €0.0028 calls (32%)
L2
Small Hosted

Gemini Flash-Lite / GPT-4o-mini (Translation & extraction)

Cost: ~€0.000112 calls (14%)
L3
Mid Hosted

Gemini Flash / Claude Haiku (Synthesis & disputes)

Cost: ~€0.0012 calls (2%)
L4
Frontier

Frontier models (Escalation only on check failure)

Cost: ~€0.010 calls (0%)

Recent AI Gateway Traces (MLflow Attached)

No live API traces recorded in current session. Mock inference simulation active.