Most LLM implementation projects don’t fail because of the technology. They fail because of decisions made before a single line of code is written: the wrong use case, the wrong architecture, the wrong success criteria, or the wrong expectations about what LLMs actually do.
This guide is for technical decision-makers and operators who want to implement LLMs in real business workflows — not build prototypes, not run proofs of concept that don’t go anywhere, but ship systems that actually work in production.
What LLMs Are Good At (And What They’re Not)
Start with a clear-eyed understanding of the capability.
LLMs are excellent at:
- Understanding natural language input with all its variation, ambiguity, and imprecision
- Generating well-structured, contextually appropriate text output
- Reasoning through multi-step problems when given appropriate context
- Extracting and structuring information from unstructured documents
- Translating between information formats (documents to JSON, bullet points to prose)
- Maintaining conversational context across multi-turn interactions
- Following complex instructions reliably when those instructions are well-specified
LLMs are unreliable at:
- Precise arithmetic and calculations (use tools for this)
- Retrieving specific factual information they weren’t trained on (use RAG or tools)
- Consistent outputs on requests where exact reproducibility is required
- Tasks requiring up-to-date information without retrieval augmentation
- Self-assessing their own uncertainty accurately in all cases
The most successful LLM implementations lean into the genuine strengths and route around the weaknesses with appropriate architecture.
The Architecture Decision That Matters Most
Before evaluating models or writing code, decide on your fundamental architecture. For business applications, the main options are:
Prompt-in, text-out (simple pipeline): An LLM receives context and a prompt, generates a response. Suitable for content generation, classification, summarisation, and extraction tasks. Simple to implement, predictable to test.
RAG (Retrieval-Augmented Generation): The LLM’s response is grounded in retrieved information from a knowledge base, database, or document store. Essential for any use case where the LLM needs to answer questions about your specific data — product catalogues, policy documents, customer records. Without RAG, you get hallucinated answers; with RAG, you get grounded answers.
Tool-using agent: The LLM can call external tools — APIs, database queries, calculators, file operations — to gather information and take actions. Necessary for any use case where the LLM needs to act, not just respond. More complex to implement and test, but required for agentic use cases.
Multi-agent system: Multiple specialised LLMs or agent instances working in coordination. A supervisor agent delegates to specialist sub-agents; results are aggregated and synthesised. Appropriate for complex, multi-domain workflows. Adds coordination overhead; don’t introduce it unless single-agent approaches are genuinely insufficient.
The right architecture is the simplest one that satisfies the requirements. Many projects over-engineer the architecture before they understand whether the fundamental approach works.
Model Selection: What Actually Differentiates Them
The LLM market has converged significantly. The leading models — GPT-4o, Claude 3.5 Sonnet/Claude 3.7, Gemini 1.5 Pro — are all capable of handling most business use cases well. The differentiation matters most at the edges:
Instruction following. For complex, multi-constraint business prompts, model ability to follow instructions precisely varies significantly. Test your actual prompts, not benchmark tasks.
Context window. Long-document use cases — contract review, lengthy correspondence analysis, large knowledge bases — require large context windows. Different models handle context length differently in practice, not just on paper.
Structured output reliability. If your system requires the LLM to generate JSON, XML, or other structured output reliably, test this specifically. Modern models with function calling / structured output modes are dramatically more reliable than models generating structured output as freetext.
Latency. For user-facing applications, latency matters. Some models are faster than others for equivalent capability. Streaming responses can improve perceived latency for longer outputs.
Cost. Token pricing varies significantly across models and providers. For high-volume applications, model economics matter. Smaller, cheaper models may be sufficient for simpler subtasks within a pipeline.
Data residency. For applications processing GDPR-regulated personal data, model deployment location matters. EU-region deployments are available from major providers, but not every model is available in every region.
The practical recommendation: evaluate two or three leading models on your actual use case with your actual prompts before committing to an architecture. Don’t extrapolate from benchmarks.
Prompt Engineering: The 80% That Determines Success
The quality of your prompts determines the quality of your outputs more than model selection. For business applications, effective prompt engineering follows consistent principles:
Be specific about the task and the output format. Vague prompts generate vague outputs. Specify exactly what you want the LLM to do, what information it should consider, and what the output should look like.
Provide examples. Few-shot prompting — providing 2-5 examples of the desired input/output pattern — dramatically improves consistency on domain-specific tasks. Don’t rely on zero-shot prompting for production systems.
Specify what to do when information is missing. LLMs will hallucinate if they don’t know how to handle gaps. Tell them explicitly: “If the information is not present in the provided context, respond with ‘Not found’ rather than inferring.”
Include constraints. “Answer only based on the provided documents. Do not use information from your training data.” This single instruction dramatically reduces hallucination in RAG systems.
Version and test your prompts. Treat prompts as code. Version control them. Test changes systematically. Document what changed and why. Prompt regression is real — a change that improves one scenario can degrade another.
Use system prompts for persistent context. The system prompt establishes the LLM’s role, capabilities, and constraints. Use it to set up the context once rather than repeating it in every user message.
Evaluation: How to Know If It’s Working
LLM evaluation is the hardest part of production deployment and the most frequently skipped. Most teams evaluate qualitatively (“it seems good”) rather than quantitatively — and then can’t reliably measure whether changes are improvements.
Build an evaluation dataset. Collect 50-200 real examples of inputs and their expected outputs from your domain experts. This becomes your test set. Run every prompt change and model update against this test set.
Define metrics appropriate to your task. For extraction tasks: precision and recall against ground truth. For classification: accuracy, F1. For generation tasks: define what “good” looks like specifically enough to measure it. Avoid metrics like “quality” that mean different things to different evaluators.
Evaluate systematically before deploying changes. Never push a prompt change to production without running it against your eval set. LLM outputs are sensitive to small prompt changes in ways that aren’t always obvious.
Implement production monitoring. Log inputs and outputs in production. Sample regularly for human review. Track downstream metrics — if your LLM powers a customer service flow, track resolution rates, escalation rates, and customer satisfaction.
RAG Architecture: Getting It Right
Retrieval-Augmented Generation is the most widely applicable architecture for business LLM applications — and the one most frequently implemented badly.
Chunking strategy matters enormously. How you split documents into chunks for retrieval determines what context the LLM can access. Too large and you waste context window; too small and you lose coherence. Hierarchical chunking (sentence → paragraph → section) with cross-reference usually outperforms flat chunking strategies.
Embedding model selection. The embedding model determines how well semantic search retrieves relevant content. Domain-specific embedding models outperform general models on specialised vocabularies. For legal, financial, or regulatory content, this matters.
Retrieval quality before generation quality. If the retrieval step doesn’t find the right content, the generation step can’t save it. Evaluate retrieval independently from generation. Track retrieval precision@k.
Query transformation. User queries are often poor retrieval queries. Techniques like HyDE (hypothetical document embeddings) and query rewriting improve retrieval performance significantly for conversational applications.
Citation and grounding. Always return source references alongside generated answers. This enables human verification, builds trust, and dramatically reduces hallucination by giving the LLM an explicit source to cite.
What Goes Wrong in Production
The failure modes that aren’t obvious until you’re live:
Prompt injection. Users discover they can include instructions in their queries that override your system prompt. Design defensive prompts and validate inputs for use cases with security implications.
Distribution shift. Your evaluation data doesn’t reflect the real distribution of production inputs. Edge cases you didn’t consider cause failures in production that didn’t appear in testing. Build in feedback mechanisms to catch and incorporate these.
Latency under load. The LLM performs fine in testing with one request at a time; it struggles under production load. Build for asynchronous processing where possible and design UX that tolerates latency.
Context window pressure. In RAG systems under real load, retrieved contexts plus conversation history can push against context window limits. This needs architectural solutions (context management, memory compression) not just more context window.
Cost at scale. Prototype costs don’t reflect production costs. A system that uses GPT-4 for every step of a complex pipeline can become extremely expensive at scale. Architect for cost from the beginning by using smaller models for simpler subtasks.
Regulatory requirements. For applications in regulated industries — financial services, healthcare, legal — automated decisions may trigger compliance requirements (explainability, human oversight, audit trails) that need to be designed in from the start, not bolted on later.
The Build vs. Buy Decision
Most organisations shouldn’t build LLM infrastructure from scratch. The vector databases, embedding pipelines, agent frameworks, and observability tools have been commoditised by the open source ecosystem and managed services.
The question is how much to build versus how much to configure and integrate. The build-vs-buy decision should be driven by:
Differentiation. Is the LLM implementation itself a competitive differentiator, or is it an operational capability that enables competitive differentiation elsewhere? Most businesses should be configuring and integrating, not building foundational infrastructure.
Domain specificity. Generic platforms often don’t handle domain-specific requirements (regulatory document formats, specialised vocabularies, sector-specific compliance) without significant customisation. When customisation is substantial, purpose-built approaches often end up being more maintainable.
Data sensitivity. Applications processing highly sensitive data may require on-premise or private cloud deployment, which constrains which managed services are viable.
Total cost of ownership. Managed services have per-unit pricing that scales with usage. Custom implementations have fixed infrastructure costs. The crossover point depends on your volume and the complexity of your requirements.
Twelve Months In: What Mature LLM Deployments Look Like
Organisations twelve months into production LLM deployments share common characteristics:
They’ve moved past single-model approaches to architectures that use different models for different tasks — smaller, cheaper models for routine classification and extraction; larger models for complex reasoning.
They have evaluation infrastructure they trust — test datasets, automated evaluation pipelines, and the ability to confidently ship or reject changes.
They’ve built feedback loops that continuously improve their systems — production logs surface failure modes, which feed into improved prompts, training data, or retrieval strategies.
They’ve learned which use cases compound — where AI automation creates compounding value as it scales — and focused investment there rather than spreading resources across too many initiatives.
The difference between organisations that get value from LLM implementation and those that struggle is almost always execution discipline, not technology.