The hard parts of shipping an LLM feature: Honeycomb's 2023 lessons, checked in 2026
In 2023 Honeycomb shipped a natural-language query assistant and published an unusually honest list of everything that was hard. Three years on, most of those problems got easier. One did not.

In May 2023, when most companies were still writing "AI strategy" slide decks, Honeycomb shipped an AI feature to every customer. Honeycomb sells observability software: engineers use it to query the data their systems produce. The new feature, Query Assistant, let people type a question in plain English, such as "which service has the highest latency?", and turned it into a real Honeycomb query.
A few weeks later Phillip Carter, the engineer behind it, published a post-mortem that has been passed around ever since: All the Hard Stuff Nobody Talks About when Building Products with LLMs. It is worth reading because it skips the demo glow and lists what actually hurt.
We build AI features for clients every week, so we went back through that list and checked every item against 2026. Here is what changed, and what did not.
The feature, in one paragraph
Query Assistant was deliberately small. It did not answer questions or chat. It produced a query the user could see, edit and run. That design choice shows up in almost every lesson below: when the model is wrong, the user sees a wrong query, not a wrong decision. The team's rule was that it is better to show something useful than nothing at all. In a 2025 look back, Honeycomb said the whole feature took about six weeks to write and ship.

The hard parts, then and now
1. The context window was too small
2023: Some customers had schemas with more than 5,000 unique fields, far more than the model could read at once. Honeycomb only sent fields that had data in the past 7 days, and truncated anything bigger. Carter called the result hit or miss.
2026: Raw size is no longer the limit. Google shipped a 2 million token window with Gemini 1.5 Pro in May 2024, and OpenAI's GPT-4.1 arrived in April 2025 with 1 million. A 5,000-field schema fits easily.
What still holds: being able to send everything does not mean you should. In our own builds, trimming the context to what matters still wins: it is cheaper, faster, and the answers are sharper. Filtering moved from a hard requirement to good engineering.
2. It was slow
2023: Each query took two to fifteen seconds or more. GPT-4 was far too slow for the job, so the team used gpt-3.5-turbo.
2026: Fast, cheap models are everywhere (OpenAI's GPT-4o mini in July 2024, Anthropic's Claude 3.5 Haiku in October 2024), and prompt caching has become standard. Anthropic launched prompt caching in August 2024 with cache reads up to 90% cheaper, and OpenAI added automatic caching at roughly half price in October 2024. Caching fits Honeycomb's pattern exactly: the same large schema goes into every request. Honeycomb's own follow-up in October 2023 reported average latency roughly halved and a much lower worst case.
3. Prompting was guesswork
2023: There was no playbook. Zero-shot prompts failed, a single example was poor, several examples worked. The popular "let's think step by step" trick made results worse on vague questions. Carter called the whole area the wild west.
2026: Two things removed most of the pain. OpenAI's Structured Outputs (August 2024) guarantee that the model returns JSON matching your schema, so broken, unparseable answers mostly disappeared. Tool calling is now mature across providers. Prompts still matter, but you no longer have to beg the model for valid output.
4. Every extra step multiplied the errors
2023: Honeycomb pointed out a simple bit of math: a step that is right 90% of the time, repeated five times, is right only about 59% of the time (0.9 to the fifth power). Chaining LLM calls gave them no real improvement.
2026: The math did not change, and it matters more now. "Agents" are exactly this: long chains of model calls. The industry answer is not a cure but discipline: fewer steps, a check after each important one, and evaluation.
5. Right versus useful
2023: Users type vague things like "slow". A strict system refuses; a helpful one guesses. Honeycomb chose to guess and let people fix the query.
2026: This trade-off is now measured instead of argued. Evaluation tools that use one model to score another's answers (Braintrust, Promptfoo, DeepEval and RAGAS, among others) turned "is it good enough?" into a number you can track release by release.
6. Prompt injection had no fix
2023: Carter compared it to SQL injection, "except worse and with no solution today". Users did try, mostly harmlessly, including attempts to pull other customers' information. Honeycomb could not prevent it, so they limited the damage: outputs were non-destructive and undoable, the model had no database access, outputs were validated, there was no open-ended chat, and there were per-user daily rate limits.
2026: Still unsolved. The OWASP Top 10 for LLM Applications 2025 (released November 2024) keeps prompt injection as a top risk, and in 2025 researchers disclosed "EchoLeak", a zero-click attack on Microsoft 365 Copilot. Honeycomb's 2023 approach, containing the model instead of trusting it, is still the right one. It is how we build: our assistants cannot touch secrets or run irreversible actions, and they ask a person before doing real work.
7. Legal work came first
2023: Honeycomb audited its model vendors (only one passed), rewrote terms of service that had not changed since 2021 within a month, disclosed the data sharing, disclaimed correctness and added an opt-out.
2026: This is now formal regulation. The EU AI Act's bans and AI literacy duties apply since February 2025, the rules for general-purpose AI models since August 2025, and most high-risk obligations from August 2026 (timeline). For companies in Indonesia, the Personal Data Protection Law (UU PDP, Law No. 27 of 2022) has been fully in force since October 2024. If customer data goes into a prompt, it is a data-processing question, not just an engineering one.
8. Early access programs fool you
2023: Honeycomb's view was that unless an early access group is large and representative, it only tells you what you want to hear. Ship to everyone.
2026: Their own numbers back it up. In the October 2023 follow-up, 82% of Enterprise and Pro Plus teams had used Query Assistant. Users were far more likely to still be writing queries by hand six weeks later (26.5% versus 4.5%), and to build complex queries (33% versus 15.7%). The model bill averaged about $30 a month, and the whole feature cost a few hundred dollars a month to run.
9. An LLM is not a product
2023: Carter predicted that thin wrappers around a completion API would disappear, and wrote that an LLM "isn't a product! It's an engine for features".
2026: That aged well. Anthropic open-sourced the Model Context Protocol in November 2024, a standard way to plug models into tools and data. In 2025 Honeycomb shipped its own hosted MCP server and an AI suite built into the product. AI became a layer inside products, not the product.
Why it matters for your business
Most of what Honeycomb found hard in 2023 is now cheap and routine: context size, speed, valid output. What remains hard is the part that was never about the model:
- Pick one narrow job. Query Assistant turned questions into queries. It did not try to be an assistant for everything.
- Contain the model. Assume it can be manipulated. Keep it away from secrets, and put a person between the model and anything irreversible.
- Measure it. Build a small set of real examples and score every change against it.
- Budget realistically. One of the first LLM features in a SaaS product ran for a few hundred dollars a month. The expensive part is the product work around it.
- Check the data rules. Under UU PDP, customer data in a prompt needs the same care as customer data anywhere else.
How we would build it today
We would ship the same feature in days, not weeks. That means a small fast model with prompt caching, structured output validated against the real schema, and an evaluation set from day one. It would have no access beyond what the feature needs, and a human would confirm anything that changes data. Then we would release it to everyone, with a switch to turn it off. The tools got better. The judgement is still the job.