Prompt engineering vs RAG vs fine-tuning
A foundation model out of the box knows a lot about the world up to its training cutoff, nothing about your company, and has its own habits about tone and format. Sooner or later every project wants it to do better: answer from the current HR handbook, write in the house style, stop inventing numbers. There are four main levers for that, and they are not interchangeable.
Generative AI exams on every cloud ask the same question in different clothes: here is a problem, which lever fixes it at the lowest cost and effort? The wrong answers are rarely absurd. They are usually a real technique applied to the wrong kind of problem, such as fine-tuning to keep up with documents that change every week. Once you can tell a knowledge problem from a behaviour problem, most of these questions answer themselves.
The four levers
Prompt engineering
You change only what you send: clearer instructions, a system prompt that sets the role and the limits, a few worked examples (few-shot prompting), a requested output format, or asking the model to reason step by step. Nothing is trained and nothing is stored.
It is the cheapest and fastest lever, and it is always the first one to try. Its limits are the context window and cost per request: every example you add is paid for on every call and adds latency. It also can't give the model facts it doesn't have. An instruction such as "answer only from current company policy" does nothing if the policy isn't in the prompt.
Retrieval-augmented generation (RAG)
At question time, the application searches your own content, picks the most relevant passages and puts them into the prompt, so the model answers from that text. Behind it is a pipeline: documents are split into chunks, each chunk is turned into an embedding (a vector), the vectors are stored in a vector index, and each question is embedded and matched against them. Better pipelines add keyword search alongside vectors (hybrid search), a reranking step, and metadata filters such as department or effective date.
RAG solves knowledge problems. Update a document, re-index it, and the next answer reflects the change without touching the model. Because the answer comes from specific passages, the application can show citations, which is often a stated requirement. It also reduces hallucination, though it doesn't eliminate it: if retrieval returns the wrong or outdated chunk, the model will faithfully answer from the wrong text. Costs are the vector store, the embedding calls, and longer prompts on every request.
Fine-tuning
You train an existing model further on your own examples, which changes its weights. In supervised fine-tuning the examples are labelled input and output pairs, typically in JSONL. Some platforms also offer preference tuning (pairs of better and worse answers), reinforcement fine-tuning (a grader or reward function scores the output), and distillation (a large teacher model generates examples to train a smaller, cheaper student).
Fine-tuning solves behaviour problems: a consistent format or schema, a house tone, domain terminology, a narrow task done more reliably, or tool calls with the right arguments. A big practical win is shorter prompts, because the dozens of few-shot examples you were paying for on every call are now built into the model. It is a poor way to add facts that change, because every change means another training run, and it gives no citations. It costs more up front: you need good labelled data, you pay for training (usually by tokens processed times epochs), and serving a custom model often carries its own hosting or storage charges.
Continued pre-training
Instead of labelled pairs, you feed the model a large volume of unlabelled domain text, such as years of legal filings or medical literature, so it absorbs the vocabulary and patterns of that domain. It sits between fine-tuning and training from scratch: much more data and compute than fine-tuning, much less than building a model. It is still not a way to track documents that change weekly, and which providers and models offer it varies, so check before assuming it is available. Training a brand new foundation model from scratch is almost never the right exam answer for a single company.
Comparing them on what exams ask about
- Cost. Prompting is cheapest to start. RAG adds a retrieval stack and longer prompts. Fine-tuning and continued pre-training add training runs and often custom model hosting. Fine-tuning can lower per-request cost later by replacing long prompts.
- Freshness. Only RAG (or grounding in live search) keeps up with changing content without retraining. Fine-tuned knowledge is frozen at the last training run.
- Hallucination. Grounding answers in retrieved sources, with citations, is the main fix for invented facts. Lowering temperature makes output more predictable, not more correct.
- Data privacy. With RAG your documents stay in your own store and can be filtered per user at query time. With fine-tuning the data shapes the weights, and a tuned model can repeat parts of its training data, so sensitive values should be removed first.
- Effort and skills. Prompting needs none. RAG needs a data pipeline. Fine-tuning needs a curated dataset and an evaluation plan to prove it beat the baseline.
How each cloud frames it
AWS. Amazon Bedrock Knowledge Bases is the managed RAG service. It ingests from data sources, chunks, embeds and stores the vectors, then retrieves passages and can return generated answers with citations. A managed knowledge base handles the whole pipeline and offers connectors such as S3, SharePoint and Confluence; a customer-managed one lets you run your own vector store, such as OpenSearch Serverless or Aurora. Changes reach it through a sync, which is incremental, so only added, changed or deleted documents are processed. Bedrock model customization covers supervised fine-tuning, reinforcement fine-tuning and distillation. On privacy, Bedrock uses your fine-tuning data only to build your private custom model, not to train the base models; it doesn't keep the training data after the job; and custom models are encrypted, with a customer managed KMS key if you want one.
Google Cloud. Google talks about grounding: connecting model output to verifiable sources, which it says reduces hallucinations, anchors answers to your data and gives links for auditability. On Gemini Enterprise Agent Platform you can ground in Google Search for current public information, or in your own data through Agent Search or the lower-level RAG Engine, where you choose the chunking, embedding model and vector store. Supervised fine-tuning is positioned for well-defined tasks with labelled data, when prompting stops giving consistent results, or to cut context length by removing few-shot examples. Google's enterprise offerings don't use customer prompts or responses to train its foundation models.
Azure. Microsoft Foundry frames fine-tuning as a way to improve accuracy on a task, align style and structure, improve tool use and reduce prompt overhead, and says plainly that it doesn't replace retrieval for current information. It also warns that fine-tuning adds training and hosting costs. RAG is built on Azure AI Search, with hybrid search, semantic ranking and document-level security trimming, and agentic retrieval that splits a complex question into several sub-queries.
Databricks. The Generative AI Engineer exam assumes RAG is the default design and tests the pipeline itself: chunk sizes that fit the embedding model, Vector Search indexes that sync from Delta tables, hybrid search, and evaluation that separates bad retrieval from bad generation.
How to choose
- Start with prompt engineering. If clear instructions and a few examples work, stop there.
- If the model lacks facts, needs private or current content, or must cite sources: use RAG.
- If the format, tone or task behaviour is still inconsistent after good prompting, and you have enough high-quality labelled examples: fine-tune.
- If the problem is both, combine them: retrieve the facts and fine-tune the behaviour.
- If the whole domain's language is unfamiliar to the model and you have a large unlabelled corpus: consider continued pre-training, where your platform offers it.
Common exam traps
- Fine-tuning on documents that change every few weeks. Every change becomes a new training run, and there are still no citations.
- A system prompt that says "use the latest policy" when the policy isn't in the prompt.
- Lowering temperature to stop hallucinations. It reduces variety, not factual errors.
- Using RAG to enforce an output format. Retrieved examples help a little, but stable formatting is a behaviour problem.
- Believing fine-tuning data trains the provider's shared base model. On the major clouds it builds a private model for you.
- Blaming the model when a RAG system cites an outdated document. That is a retrieval or data problem: fix the index, filters or source data.
A worked example
An insurance company's claims assistant has two complaints. Adjusters say it quotes coverage limits from last year's policy wording, which is revised every quarter. Managers say its claim summaries ignore the required five-heading layout, even though the prompt now includes fifteen example summaries and has become slow and expensive. The team has thousands of summaries that reviewers have already approved.
These are two different problems. The outdated limits are a knowledge problem: the wording changes quarterly and adjusters want to see the source, so the policy documents belong in a RAG index that is re-synced when they change. The layout is a behaviour problem that prompting has already failed to fix, and there is a large set of approved examples, so supervised fine-tuning on those summaries fits. It also lets the team drop the fifteen examples from every prompt, which fixes the cost and latency. Fine-tuning on the policy wording would solve neither problem well, and continued pre-training is far more than the situation needs.
Practice questions
- An HR assistant for policies that change every few weeks
- A summarizer that needs 25 examples in every prompt
- What happens to data used to customize a Bedrock model
- An assistant citing a replaced parental-leave policy
- Prompt templates that still drift on long contracts
- A drafting assistant with two different problems