AI Engineering

Language model features that survive the next prompt change.

The hard part of shipping with a language model is knowing it'll still behave after the next prompt edit, proving it in CI, keeping untrusted text from becoming instructions, and being able to say what it costs. Digital Lotus builds LLM features with the same discipline you would expect around any other production dependency.

2-tierDeterministic assertions block the merge; model-graded scores stay a trend line
1 schemaDrives the database, the API contract, the UI types, the tool call, and the tests
0Model output reaching your data unverified against your own rules
Every callTokens, spend, latency, retries and fallback outcomes, on a dashboard

How it's built

Maintainable LLM features.

Proof It Still Works

A prompt change is a code change, and it deserves the same gate. Fixtures run repeatedly with deterministic assertions blocking a merge at explicit pass rates, and model-graded scores kept as a trend line rather than a gate. Prompts live in the repo as versioned, content-hashed files, so every output records which version produced it.

Output You Can Rely On

One schema definition drives the database, the API contract, the UI types, the tool call handed to the model, and the tests. A forced tool call leaves no output path that isn't the schema, and a validation failure feeds the specific fields back rather than asking for a rewrite.

Checked Against Your Real Rules

Treat the model as an untrusted planner and check its work against a verifier built from the same functions your real logic uses. Every violation is returned at once so a rejection costs one round trip, and continued failure falls back to a deterministic path that records that it did.

Grounded In Your Data

When your content is too large to paste into a prompt and only some of it matters to any one request, retrieval is the answer: chunk it, embed it, and put only the relevant passages in front of the model. Knowing when it is not the answer matters just as much.

Safe At The Input Boundary

Any text a user controls that reaches a system prompt is an injection path. Red-team fixtures embed injection attempts in the fields the model reads as data, asserted separately so a bypass cannot hide inside an aggregate. Model proposals are held until a person accepts them, so nothing a model returns reaches your database on its own.

Visible In Production

Spans per generation covering input and output tokens, computed spend, latency, retries, and fallback outcomes, exported and put on a dashboard. Recorded fixtures let a stakeholder click through a working feature with no live spend, and make it reproducible under test.

New to some of this vocabulary?

Evals, RAG, tool use, prompt injection — the terms above in plain English, including what it is, what problem it solves, and when it's worth the complexity.

Go to the AI terms

AI Project Demo

Project showcasing AI integration and features.

Recipe Generator

  • AI
  • Full-stack
  • Next.js
  • TypeScript
  • tRPC
  • Drizzle
  • SQLite
  • Anthropic API
  • Zod schemas
  • Vitest evals
  • LLM-as-judge
  • OpenTelemetry
  • Prometheus / Grafana

A meal planner that turns body metrics into macro targets and schedules a week of dinners around cook days and leftovers. Claude proposes the plan and writes the recipes; a verifier built from the same functions the deterministic planner uses checks every result in code, returns every violation at once, and falls back to deterministic planning when the model cannot satisfy the constraints.

AI features used

  • Agentic planner given four tools, then treated as untrusted and checked by a verifier built from the deterministic planner's own functions
  • Forced-tool-call generation validated against one shared zod schema, retrying with the specific failed fields rather than asking for a rewrite
  • Two-tier eval suite in CI: deterministic assertions gate the merge, a model judge tracks the trend without ever blocking it
  • Red-team fixtures embedding prompt injection in the fields the model reads as data, asserted separately so a bypass cannot hide in an aggregate
  • Versioned, content-hashed prompts, with the hash and model string recorded on every generation
  • Local MiniLM embeddings in sqlite-vec for natural-language search and exemplar retrieval for the generator
  • A setup interview whose proposals are held in local state until a person accepts them
  • OpenTelemetry spans per generation for tokens, computed spend, latency, retries, and fallback outcomes
  • Recorded-fixture demo mode, so the public demo is reproducible and costs nothing to run
The week planner showing each day's macro targets and planned dinner
The grocery list, grouped by aisle, with each ingredient traced back to its recipe

Background

In progress

M.S. Artificial Intelligence

Johns Hopkins University.

Since 2016

Production systems first

A decade of full-stack delivery across scoping, architecture, implementation, and technical leadership.

Teaching

Adoption is a people problem

Expertise in mentoring developers and building internal learning programs. Getting a team genuinely productive with AI tooling requires hands-on guidance as well as the right workflows.

How teams start

Pick the one that fits where you are.

01

Readiness review

You have an LLM feature in production or near it, and no way to tell whether a prompt change made it better or worse. We assess what exists, then stand up evals, structured output, and the metrics that make the next change measurable.

02

Build the feature

An LLM feature designed the way the rest of your system is designed: schema-driven, verified against real rules, observable, with a deterministic path for when the model cannot deliver.

03

Get the team fluent

Tooling setup, working practices, and responsible-use guidance for a team adopting AI. Delivered as mentoring rather than a slide deck, and ending with your developers empowered to do this autonomously.

Let's make things happen

Let's chat and find the right plan for your needs.