The moving parts of building with large language models, in plain terms: what the words mean, what an API call looks like, how RAG and tools fit together, how to test, and the security risks to design for. Nothing here is tied to one vendor.
The words
| Term | What it means |
|---|---|
| Model | The trained network you send text (or images) to and get text back from. Different models trade quality, speed and cost. |
| Token | The unit a model reads and writes: a word, part of a word or a punctuation mark. Limits and prices are counted in tokens. |
| Context window | How many tokens one request can hold: your instructions, the conversation so far, any documents, and the answer. |
| System prompt | Standing instructions for the model, set by the application rather than typed by the user. |
| Temperature | How much randomness goes into picking each token. Low for factual or structured output, higher for drafting. |
| Embedding | A list of numbers that represents the meaning of a piece of text, so similar texts end up close together. |
| Vector database | A store that finds the embeddings closest to a query: the search half of RAG. |
| RAG | Retrieval-augmented generation: look up relevant documents first and put them in the prompt, so the answer comes from your data. |
| Fine-tuning | Training a model further on your own examples. Changes its style or skills; it is not the way to give it facts that change. |
| Tool use | The model asks your code to run a function (search, look up a ticket, run a query) and gets the result back. |
| Agent | A loop where the model plans, calls tools, reads the results and decides the next step until the task is done. |
| Hallucination | A confident answer that is wrong or made up. Every design choice below exists partly to catch it. |
| Eval | A repeatable test of the model’s answers against cases you know the answer to. |
One API call
Most model APIs take a JSON request like this one and return JSON with the answer and a count of the tokens used. Field names differ between providers (some send the system prompt as a message with the role system), so check the provider’s reference.
{
"model": "the-model-id",
"system": "You answer IT support questions. Use only the documents provided. If they do not cover it, say so.",
"messages": [
{ "role": "user", "content": "Why does Outlook keep asking for my password?" }
],
"max_tokens": 500,
"temperature": 0
}
- You pay for input and output tokens, so a long conversation or a big document costs more on every call.
- Keep the API key in an environment variable or a secret store, never in code, a repository or the prompt.
- Rate limits answer with HTTP 429. Retry after a pause that grows each time (exponential backoff), not in a tight loop.
- As a rough guide, a token is about four characters of English text. It varies by model and language, so measure with the provider’s own counter.
Prompts that work
- Say who the answer is for and what it is for.
- Give the facts it needs, rather than hoping it knows them.
- Show one or two examples of a good answer.
- Ask for a fixed format (a list, a table, JSON with named fields) when code will read the result.
- Tell it what to do when it does not know: say so, and do not guess.
- Keep instructions and data apart, for example by marking where a pasted document starts and ends. This helps, but it does not stop prompt injection on its own.
RAG in five steps
- Split your documents into chunks of a few paragraphs.
- Turn each chunk into an embedding and store it with a link back to its source.
- Turn the question into an embedding and fetch the closest chunks, only from documents this user is allowed to read.
- Put those chunks in the prompt, with an instruction to answer only from them and cite which one.
- Show the sources with the answer, so a person can check it.
Tools and agents
- The model only asks for a tool call. Your code decides whether to run it, runs it, and sends back the result.
- Give each tool the least access that does the job: a read-only account for lookups, a separate one for changes.
- Ask a person to confirm anything that deletes, sends, pays or changes access.
- Log every tool call with its inputs and results.
- Cap the number of steps and the spend per task, so a loop cannot run for ever.
Test before you ship
- Collect real questions with known good answers: that is your eval set. Twenty good cases beat none.
- Check what can be checked by code: valid JSON, the right fields, a pattern match, an exact value.
- Have a person review a sample of the rest.
- Run the whole set again after any change of prompt, model or data, and compare with the last run.
- Output varies from run to run. A temperature of 0 reduces the variation but does not guarantee the same answer every time.
Security: the OWASP Top 10 for LLM applications (2025)
| ID | Risk | What it means for you |
|---|---|---|
| LLM01 | Prompt Injection | Text in an email, web page or document can carry instructions the model follows. Treat all content as untrusted. |
| LLM02 | Sensitive Information Disclosure | The model can repeat secrets or personal data it was given. Do not put in what the user may not see. |
| LLM03 | Supply Chain | Models, plugins and datasets are dependencies: check where they come from, like any package. |
| LLM04 | Data and Model Poisoning | Tampered training, fine-tuning or retrieval data changes the answers. Control what goes into the index. |
| LLM05 | Improper Output Handling | Model output is untrusted input to the next system. Validate it before it reaches a shell, a query or a web page. |
| LLM06 | Excessive Agency | Tools with more permissions than the task needs. Least privilege, and a human confirms anything destructive. |
| LLM07 | System Prompt Leakage | Assume the system prompt can be extracted. Never put keys, passwords or access rules in it. |
| LLM08 | Vector and Embedding Weaknesses | A shared index can leak documents across users. Filter retrieval by the user’s own permissions. |
| LLM09 | Misinformation | Plausible but wrong answers. Cite sources, check facts, and say what was not verified. |
| LLM10 | Unbounded Consumption | Runaway loops and huge requests cost money and can take a service down. Cap tokens, steps and spend. |
Before it goes live
- Where does the data in each prompt come from, and who is allowed to see it?
- Where do the prompts and answers get stored, for how long, and does the provider train on them?
- What is the worst thing a tool can do if the model is tricked into calling it?
- Who reviews the answers, and how does a user report a bad one?
- What does it cost per request at the expected volume, and where is the cap?