AI Terms
The vocabulary, in plain English — including what it is, what problem it solves, and when it's worth the complexity.
For how these ideas fit together in a real build, see AI engineering.
Discovery phase
The answers worth knowing before starting a project.
- AI safety
- The work of making a model behave as intended and refuse to cause unintentional harm. Distinct from security, which is about people attacking your system — safety is about the model itself. Most of it is done by the provider before you ever see the model, and the rest is decided by what you let it do, and what you check before it does it.
- Alignment
- Whether a model's behaviour matches what you meant, rather than what you literally said. You notice it when the model follows your instruction and defeats your purpose. It's why a system prompt gets written, tested, and versioned.
- Constitutional AI
- Training a model against a written set of principles instead of case-by-case human ratings. Anthropic publishes Claude's constitution. It explains why a model refuses some things and hedges on others, and a published set of principles is one you can read and disagree with before building on it.
- Bias
- Systematic skew inherited from training data and reproduced in output. If a feature touches hiring, lending, housing or insurance, it's legal exposure as well as an ethical problem. It has to be tested on your data and your population. See: Massachusetts tenancy AI discrimination lawsuit, or Workday AI screening bias case.
- Explainability (XAI)
- Whether anyone can say why a model produced a particular output to ensure transparency, meet regulatory requirements, and build trust. Highly regulated sectors may require explainability.
- Disclosure
- Telling people when they're talking to a model, or when what they're reading was generated by one. Increasingly a legal requirement rather than a courtesy, and a product decision regardless since users who work it out for themselves later trust the rest of the product less.
The basics
Basic terms for understanding key concepts in AI and how they are applied.
- Large language model (LLM)
- A model trained to predict the next piece of text. Everything it does — answering, summarising, writing code — is that one trick applied well. It has no memory between calls and no access to your systems unless explicitly granted.
- Token
- The unit models read and bill in. You pay per token in and per token out, which is why prompt size is a cost decision, not just a design one.
- Context window
- How much the model can hold at once, including your instructions, the conversation, and any data you paste in. Large but finite, and filling it with marginally relevant text makes answers worse as well as more expensive.
- System prompt
- The standing instructions the model gets before the user's message. It is where the rules live, and it should be treated as source code: versioned, reviewed, and tested when it changes.
- Temperature
- How much randomness to allow. Low for extraction and classification, higher for drafting. It reduces variation but never removes it.
- Hallucination
- Fluent, confident, wrong. This is what next-token prediction does when it lacks facts. The fix is architectural: give it better data, or check the output.
- Reasoning model
- A model trained to spend extra tokens working through a problem before it answers. It buys accuracy on multi-step work, and it pays for that in latency and cost. It buys very little on the extraction and classification for most features, so it's worth discussing before accepting the premium.
- Inference
- One run of a model over one input. Billing is usually per input and output token, while latency is per inference. Each inference adds network plus model time, so more calls usually means slower features and adds more possible failure points in the form of potential timeouts.
Getting usable output
What turns a chat box into a real feature is output with a shape you can rely on.
- Structured output
- Making the model return data in a fixed shape. Without it you are writing parsers against text that changes whenever the prompt does.
- Tool use (function calling)
- Handing the model a set of functions it may call, with typed arguments. Forcing a tool call is also the cleanest way to guarantee structured output: the model has no path that isn't the schema.
- Schema
- The contract the output must satisfy. The valuable version is defined once and reused everywhere — database, API, UI types, the tool definition, and the tests — so the shape doesn't drift between them.
- Validation and repair
- Checking output against the schema and, on failure, telling the model exactly which field was wrong instead of asking for a rewrite. Cheaper, faster, and far more likely to succeed on the retry.
- Few-shot prompting
- Showing the model two or three worked examples instead of describing what you want in prose. It is the cheapest quality improvement available and it routinely beats a longer instruction — which is most of “prompt engineering”.
- Streaming
- Sending the answer token by token as it is produced rather than waiting for all of it. It doesn't make generation faster. It makes the wait visible, which keeps the feature from feeling broken.
Grounding it in the data
How a general model comes to know things specific to the data.
- Embedding
- A list of numbers representing the meaning for a chunk of text. Similar meanings land close together, which is what lets you search by sense rather than by keyword.
- Vector database
- Storage optimised for finding the nearest embeddings to a query. At a small scale a plain database column does the job, so the index earns its keep with volume.
- Retrieval-augmented generation (RAG)
- Search your own content for the passages relevant to a request, then put only those in the prompt. Exists because context windows are finite and the model does not know your data and is the right tool when your content is too large to inline and only variably relevant.
- Knowledge cutoff
- The date after which the model was trained. It's why a model can be confidently wrong about recent events and confidently ignorant of your product in its current state. Retrieval is the fix.
- Chunking
- Splitting documents into retrievable pieces. Unglamorous and decisive. If you chunk badly the retrieval returns fragments that read as nonsense, no matter how good the model is.
- Fine-tuning
- Further training on your own examples. It teaches form and tone well, and facts poorly, so if the goal is for the model to know your data, retrieval is almost always the cheaper and more maintainable answer.
- Hybrid search
- Running keyword search and vector search together and combining the results. Vector search alone misses exact strings because there may be nothing semantically distinctive about them. A keyword search plus a vector search returns exact hits plus the related context.
- Reranking
- A second, slower pass that re-scores the passages retrieval returned and keeps the best few. Retrieval is optimised for finding plausible candidates quickly while reranking is optimised for putting the right one first.
- Data readiness
- Whether the content you want a model to use is in a state that can be properly consumed. This gets assessed first, or it gets discovered late and expensively.
Knowing whether it works
The step that turns a vague 'it's not working' into something you can actually fix.
- Eval
- An automated test for model output. The distinguishing feature is that it runs repeatedly against the same input, because a constraint that holds once may not hold every time. What you track is the pass rate.
- Golden set
- The fixed collection of cases every change is measured against. The good ones grow out of real failures rather than cases someone first imagined.
- LLM-as-judge
- Using a model to score output that cannot be checked mechanically like tone, coherence, and helpfulness. Useful as a trend line. Dangerous as a merge gate, because a judge is a model too, and gating on one makes your pipeline as unpredictable as the judge.
- Verifier
- Code that checks model output against your real rules before anything is accepted. The important property is that it should reuse the same functions the rest of your system uses. A verifier with its own copy of the logic is just a second thing to get wrong.
- Regression
- Something that used to work and now doesn't. Common after a prompt edit, and invisible without evals, which is why a prompt change with no eval run is a deploy with no test suite.
- Drift
- The same prompt producing different results over time because the provider changed the model underneath it. Nothing in your code changed, so nothing in your test suite ran. It's the reason evals have to run on a schedule and not only on a merge.
- Deterministic fallback
- The non-AI path you take when the model fails, times out, or cannot satisfy the constraints. If a feature has no acceptable answer to “what happens when the model is down”, it is not finished.
Risk
The failure modes that matter once real users and real data are involved.
- Prompt injection
- Text the model reads as data containing instructions it follows anyway, in a support ticket, a CV, a web page, or a filename. It is the defining security problem of LLM features, and it's not solved by asking the model nicely to ignore it.
- Indirect injection
- The same attack arriving through content your system fetched rather than something the user typed. Harder to spot, because nobody in the conversation ever wrote the malicious text.
- Data leakage
- Sensitive information reaching somewhere it shouldn't, such as another user's answer, a log, or a third-party API. Worth deciding deliberately what leaves your infrastructure, since some models can run locally and never make a network call.
- Training on your data
- Whether the provider uses what you send to improve their models. Consumer and business tiers of the same product often differ on this and the default may be opt-out instead of opt-in.
- Guardrail
- A check around the model rather than inside it. Rules should be validated in code, instead of another prompt asking a model to behave.
- Jailbreak
- A user talking the model out of its own instructions, in their own session, deliberately. Distinct from prompt injection, where the attacker is text your system read from somewhere else and the victim is your user. Worth separating because the damage differs.
- Human in the loop
- A human approval before anything consequential happens. The cheapest real safeguard there is, and the right default anywhere a model's output would otherwise write to a database, send a message, or move money.
- Failure mode
- How a feature behaves when it's wrong, which for a model is quietly rather than loudly. A model returns a fluent, plausible, incorrect answer and no warning. Deciding what should happen in that case is a product decision.
Running it in production
What changes once the feature is live.
- Prompt caching
- Reusing an unchanged chunk of prompt across calls at a reduced rate. Meaningful savings when a long system prompt is sent on every request, which is most production features.
- Cost per request
- Total tokens in and out for one user action, priced at the model’s current rate, given in a range. Think of this like cloud spend: it can be forecast from assumptions, but it's managed in practice with telemetry once real usage starts.
- Latency and time to first token
- Total time versus time until something appears. Streaming improves the second without changing the first, which is often enough for a good user experience.
- Model pinning
- Naming an exact model version. It turns drift from a model update from something that happens to you into something you schedule.
- Agent
- This term is loosely used, and worth pinning down. Usefully: a model given tools and allowed to decide which to call, in a loop, until a goal is met. The engineering question isn't how autonomous it is, but instead what checks the work, and what happens when the work is wrong.
- Observability
- Recording what each call cost, how long it took, how often it retried, and how often it fell back. Without it, “the AI isn't working” is a suspicion nobody can confirm or refute.