Three optimization patterns we apply on most production LLM engagements: semantic routing, custom calibration, prompt compression. Each addresses a different bottleneck. Stacked, they typically deliver 5-10× cost reduction without quality loss.
Half of the engagements we turn down come from teams trying to use an LLM where a simpler tool would have shipped in a week. Here's the triage we run before agreeing to take a project.
Most production LLM systems track latency and cost. They don't track the things that matter for ongoing operations. Here are the three measurements we insist on before declaring a deployment ready.
API pricing is per-token and scales linearly. On-prem cost is mostly fixed plus operational overhead. The honest TCO at three usage tiers, with hidden costs disclosed.
Three constraints decide whether a legal-tech LLM deployment can happen on a third-party API: privilege, retention, and audit. Self-hosting clears all three; APIs clear about one and a half.
Financial services has three constraints LLM deployments collide with: regulatory examination, supervisory recordkeeping, and customer data residency. Self-hosting addresses all three; the GPU bill is the easy part.
Most enterprise data residency requirements were written for SaaS and don't translate cleanly to LLM APIs. On-prem deployment sidesteps the translation problem. Here's the analysis we run on engagements.
Most AI initiatives go through the same three-way decision: buy a SaaS product, build an in-house solution, or hire a firm to do it. Each is right in different contexts. Here's how we triage.
When an auditor asks 'show me what the AI did,' the answer needs to be specific, complete, and reproducible. Most production LLM systems can't deliver one of those three. Here's the architecture that does.
Mid-market SaaS company, support-deflection chatbot, $34K/month OpenAI bill. Six weeks of work, prompts down 70%, accuracy flat. Bill dropped to $11K/month.
Regional law firm using OpenAI for document classification and clause extraction at $42K/month. Migrated to on-prem Llama-3.3 70B with custom calibration. New monthly bill: $4,200 amortized, paid back in 9 months.
Subscribe
Get new deep-dives in your inbox
Roughly one article a fortnight. Technical, no marketing fluff, unsubscribe any time. We never share your address.