The first time a developer adds an LLM call to an app, their test suite has a small crisis. The function they want to test returns different output every time, on purpose. Assert on the response text and the test is flaky by design; don't test it at all and the most important feature in the app ships on vibes.
After building several LLM-powered Flutter apps, I've landed on a simple reframe that dissolves most of the problem:
You are not testing the model. You are testing everything around the model.
The model's quality is a product concern you evaluate; your app is a deterministic machine that assembles requests, parses responses, and renders states. All of that is testable with completely ordinary techniques — once you draw the boundary in the right place.
Draw the boundary: one interface, typed both ways
Every LLM interaction in my apps goes through a single Dart interface with typed inputs and typed outputs:
abstract class PlannerService {
Future<Either<Failure, DayPlan>> generatePlan(PlanContext context);
}
No widget, controller, or repository ever sees a prompt string or raw completion. That one decision creates the testing seam: in tests, PlannerService is swapped for a fake that returns whatever you want — instantly, deterministically, offline.
Layer 1 — Unit tests with fixture responses
Everything downstream of the model is tested against fixtures: real captured responses, saved as JSON files, replayed forever.
- The parser gets the widest fixture set: perfect responses, responses with missing fields, wrong enum values, truncated JSON, empty lists, the model wrapping JSON in markdown fences (it will), and plain refusal text. Every weird response from real usage gets captured and added — the fixture folder becomes a regression archive of every way the model has ever surprised you.
- Controllers get fakes returning success, each
Failuretype, and slow futures — asserting the state machine walks loading → data / error correctly.
None of this touches the network. It runs in milliseconds in CI, and it covers the code where LLM apps actually break — because in practice, the bug is almost never "the model said something wrong." It's "the app handled the model's output wrong."
Layer 2 — Contract tests on the request side
The other half people forget: testing what you send. Context assembly is real logic — it merges user state, history, and settings into a request — and when it's wrong, output quality drops silently. No crash, no error, just a model working with bad information. These bugs are invisible in production, which is exactly why they need tests:
test('plan request includes today\'s completed items only', () {
final ctx = buildPlanContext(user: fixtureUser, history: fixtureWeek);
expect(ctx.recentActivity, hasLength(3));
expect(ctx.recentActivity, everyElement(isA<CompletedItem>()));
});
Assert on the structured context object, not on the final prompt string — prompt wording changes constantly, and tests pinned to phrasing die within a week.
Layer 3 — Schema validation at runtime, tested in CI
Because the model returns structured data (if yours doesn't, start there), the schema validator is itself a unit under test. The rule I enforce: nothing model-generated reaches a widget without passing validation. The validator's tests define your app's actual contract with the model — required fields, value ranges, list bounds, string sanitisation. When a model upgrade subtly changes output shape, this layer is what notices, loudly, in staging rather than in a user's hands.
Golden tests: for states, not sentences
Golden (screenshot) tests feel useless with non-deterministic content — until you feed them fixtures. Freeze the input, and the rendering is deterministic. I keep goldens for each state of an AI surface: streaming-in-progress, rendered plan, partial data, each error variant, the regenerate affordance, offline-with-cache. The goldens aren't checking the model's words; they're checking that a 2,000-token response doesn't blow up the layout and that error states actually exist visually. Long-content fixtures earn their keep here — real model output is reliably longer than designer lorem ipsum.
And the model itself?
Model output quality genuinely can't be unit-tested, but it can be evaluated: a small suite of representative context fixtures run against the live model on demand — before a model swap or prompt change, not on every CI run — with cheap automatic checks (valid schema? right length? no vendor-specific artifacts?) and a human eyeball on a sample. Keep this out of the deploy path; it's a dashboard, not a gate.
The takeaway
Split the problem and none of it is scary:
| Concern | Technique | Deterministic? |
|---|---|---|
| Request assembly | Unit tests on context objects | Yes |
| Response handling | Fixture-driven parser tests | Yes |
| State machine / UX | Fakes behind the service interface | Yes |
| Layout under real output | Golden tests on fixtures | Yes |
| Model quality | Offline eval suite, human-reviewed | No — and that's fine |
The non-determinism never actually enters your test suite. It stays where it belongs — behind the interface, evaluated on your schedule, in an app engineered to survive whatever comes back.