AI Terms

The vocabulary, in plain English — including what it is, what problem it solves, and when it's worth the complexity.

For how these ideas fit together in a real build, see AI engineering.

Discovery phase

The answers worth knowing before starting a project.

AI safety
The work of making a model behave as intended and refuse to cause unintentional harm. Distinct from security, which is about people attacking your system — safety is about the model itself. Most of it is done by the provider before you ever see the model, and the rest is decided by what you let it do, and what you check before it does it.
Alignment
Whether a model's behaviour matches what you meant, rather than what you literally said. You notice it when the model follows your instruction and defeats your purpose. It's why a system prompt gets written, tested, and versioned.
Constitutional AI
Training a model against a written set of principles instead of case-by-case human ratings. Anthropic publishes Claude's constitution. It explains why a model refuses some things and hedges on others, and a published set of principles is one you can read and disagree with before building on it.
Bias
Systematic skew inherited from training data and reproduced in output. If a feature touches hiring, lending, housing or insurance, it's legal exposure as well as an ethical problem. It has to be tested on your data and your population. See: Massachusetts tenancy AI discrimination lawsuit, or Workday AI screening bias case.
Explainability (XAI)
Whether anyone can say why a model produced a particular output to ensure transparency, meet regulatory requirements, and build trust. Highly regulated sectors may require explainability.
Disclosure
Telling people when they're talking to a model, or when what they're reading was generated by one. Increasingly a legal requirement rather than a courtesy, and a product decision regardless since users who work it out for themselves later trust the rest of the product less.

The basics

Basic terms for understanding key concepts in AI and how they are applied.

Large language model (LLM)
A model trained to predict the next piece of text. Everything it does — answering, summarising, writing code — is that one trick applied well. It has no memory between calls and no access to your systems unless explicitly granted.
Token
The unit models read and bill in. You pay per token in and per token out, which is why prompt size is a cost decision, not just a design one.
Context window
How much the model can hold at once, including your instructions, the conversation, and any data you paste in. Large but finite, and filling it with marginally relevant text makes answers worse as well as more expensive.
System prompt
The standing instructions the model gets before the user's message. It is where the rules live, and it should be treated as source code: versioned, reviewed, and tested when it changes.
Temperature
How much randomness to allow. Low for extraction and classification, higher for drafting. It reduces variation but never removes it.
Hallucination
Fluent, confident, wrong. This is what next-token prediction does when it lacks facts. The fix is architectural: give it better data, or check the output.
Reasoning model
A model trained to spend extra tokens working through a problem before it answers. It buys accuracy on multi-step work, and it pays for that in latency and cost. It buys very little on the extraction and classification for most features, so it's worth discussing before accepting the premium.
Inference
One run of a model over one input. Billing is usually per input and output token, while latency is per inference. Each inference adds network plus model time, so more calls usually means slower features and adds more possible failure points in the form of potential timeouts.

Getting usable output

What turns a chat box into a real feature is output with a shape you can rely on.

Structured output
Making the model return data in a fixed shape. Without it you are writing parsers against text that changes whenever the prompt does.
Tool use (function calling)
Handing the model a set of functions it may call, with typed arguments. Forcing a tool call is also the cleanest way to guarantee structured output: the model has no path that isn't the schema.
Schema
The contract the output must satisfy. The valuable version is defined once and reused everywhere — database, API, UI types, the tool definition, and the tests — so the shape doesn't drift between them.
Validation and repair
Checking output against the schema and, on failure, telling the model exactly which field was wrong instead of asking for a rewrite. Cheaper, faster, and far more likely to succeed on the retry.
Few-shot prompting
Showing the model two or three worked examples instead of describing what you want in prose. It is the cheapest quality improvement available and it routinely beats a longer instruction — which is most of “prompt engineering”.
Streaming
Sending the answer token by token as it is produced rather than waiting for all of it. It doesn't make generation faster. It makes the wait visible, which keeps the feature from feeling broken.

Grounding it in the data

How a general model comes to know things specific to the data.

Embedding
A list of numbers representing the meaning for a chunk of text. Similar meanings land close together, which is what lets you search by sense rather than by keyword.
Vector database
Storage optimised for finding the nearest embeddings to a query. At a small scale a plain database column does the job, so the index earns its keep with volume.
Retrieval-augmented generation (RAG)
Search your own content for the passages relevant to a request, then put only those in the prompt. Exists because context windows are finite and the model does not know your data and is the right tool when your content is too large to inline and only variably relevant.
Knowledge cutoff
The date after which the model was trained. It's why a model can be confidently wrong about recent events and confidently ignorant of your product in its current state. Retrieval is the fix.
Chunking
Splitting documents into retrievable pieces. Unglamorous and decisive. If you chunk badly the retrieval returns fragments that read as nonsense, no matter how good the model is.
Fine-tuning
Further training on your own examples. It teaches form and tone well, and facts poorly, so if the goal is for the model to know your data, retrieval is almost always the cheaper and more maintainable answer.
Reranking
A second, slower pass that re-scores the passages retrieval returned and keeps the best few. Retrieval is optimised for finding plausible candidates quickly while reranking is optimised for putting the right one first.
Data readiness
Whether the content you want a model to use is in a state that can be properly consumed. This gets assessed first, or it gets discovered late and expensively.

Knowing whether it works

The step that turns a vague 'it's not working' into something you can actually fix.

Eval
An automated test for model output. The distinguishing feature is that it runs repeatedly against the same input, because a constraint that holds once may not hold every time. What you track is the pass rate.
Golden set
The fixed collection of cases every change is measured against. The good ones grow out of real failures rather than cases someone first imagined.
LLM-as-judge
Using a model to score output that cannot be checked mechanically like tone, coherence, and helpfulness. Useful as a trend line. Dangerous as a merge gate, because a judge is a model too, and gating on one makes your pipeline as unpredictable as the judge.
Verifier
Code that checks model output against your real rules before anything is accepted. The important property is that it should reuse the same functions the rest of your system uses. A verifier with its own copy of the logic is just a second thing to get wrong.
Regression
Something that used to work and now doesn't. Common after a prompt edit, and invisible without evals, which is why a prompt change with no eval run is a deploy with no test suite.
Drift
The same prompt producing different results over time because the provider changed the model underneath it. Nothing in your code changed, so nothing in your test suite ran. It's the reason evals have to run on a schedule and not only on a merge.
Deterministic fallback
The non-AI path you take when the model fails, times out, or cannot satisfy the constraints. If a feature has no acceptable answer to “what happens when the model is down”, it is not finished.

Risk

The failure modes that matter once real users and real data are involved.

Prompt injection
Text the model reads as data containing instructions it follows anyway, in a support ticket, a CV, a web page, or a filename. It is the defining security problem of LLM features, and it's not solved by asking the model nicely to ignore it.
Indirect injection
The same attack arriving through content your system fetched rather than something the user typed. Harder to spot, because nobody in the conversation ever wrote the malicious text.
Data leakage
Sensitive information reaching somewhere it shouldn't, such as another user's answer, a log, or a third-party API. Worth deciding deliberately what leaves your infrastructure, since some models can run locally and never make a network call.
Training on your data
Whether the provider uses what you send to improve their models. Consumer and business tiers of the same product often differ on this and the default may be opt-out instead of opt-in.
Guardrail
A check around the model rather than inside it. Rules should be validated in code, instead of another prompt asking a model to behave.
Jailbreak
A user talking the model out of its own instructions, in their own session, deliberately. Distinct from prompt injection, where the attacker is text your system read from somewhere else and the victim is your user. Worth separating because the damage differs.
Human in the loop
A human approval before anything consequential happens. The cheapest real safeguard there is, and the right default anywhere a model's output would otherwise write to a database, send a message, or move money.
Failure mode
How a feature behaves when it's wrong, which for a model is quietly rather than loudly. A model returns a fluent, plausible, incorrect answer and no warning. Deciding what should happen in that case is a product decision.

Running it in production

What changes once the feature is live.

Prompt caching
Reusing an unchanged chunk of prompt across calls at a reduced rate. Meaningful savings when a long system prompt is sent on every request, which is most production features.
Cost per request
Total tokens in and out for one user action, priced at the model’s current rate, given in a range. Think of this like cloud spend: it can be forecast from assumptions, but it's managed in practice with telemetry once real usage starts.
Latency and time to first token
Total time versus time until something appears. Streaming improves the second without changing the first, which is often enough for a good user experience.
Model pinning
Naming an exact model version. It turns drift from a model update from something that happens to you into something you schedule.
Agent
This term is loosely used, and worth pinning down. Usefully: a model given tools and allowed to decide which to call, in a loop, until a goal is met. The engineering question isn't how autonomous it is, but instead what checks the work, and what happens when the work is wrong.
Observability
Recording what each call cost, how long it took, how often it retried, and how often it fell back. Without it, “the AI isn't working” is a suspicion nobody can confirm or refute.

Let's make things happen

Let's chat and find the right plan for your needs.