AI · Jun 2026 · 9 min read

Adding real AI features — not a chatbot bolted on

How I wire LLMs into Flutter apps so the AI is core to the product — not a gimmick users skip after day two. Lessons from building Tertio's daily coach.

Every founder brief I've received in the last two years has included some version of the same line: "and we want AI in it." Fair enough — but when I ask what the AI should do, the answer is usually "a chatbot, maybe?"

That instinct produces the most skippable feature in modern apps: a floating chat button that opens a generic assistant, gets tried twice, and never gets tapped again. I've built AI-first products — including Tertio, a wellness app where an LLM plans your entire day — and the difference between AI that retains users and AI that decorates a settings menu comes down to a few decisions you make before writing any code.

The bolted-on chatbot test

Here's the test I run on every AI feature idea: if the model gave a mediocre answer, would the product still work?

If yes, the AI is decoration. A chatbot bolted onto a habit tracker can be wrong, slow, or bland — the habit tracker underneath still functions, which means users learn to ignore the chatbot. Nothing in the product's core loop depends on it.

In Tertio, the answer is no. The core loop is the AI: you tell it how you slept, what your energy is like, and what today looks like — it produces your day plan. If the model output were mediocre, the product would be mediocre. That sounds risky, and it is — but it's also the only configuration where the AI earns its place. Users open the app for the AI output, not despite it.

Structure the output, not the conversation

The single biggest technical decision: the LLM should return data, not prose.

A chat interface makes the model's words the UI. That caps your design at "a wall of text with an avatar." Instead, I define a strict response schema — for Tertio's planner, that's a list of plan blocks with times, categories, durations, and a short rationale — and the model must fill that structure. The Flutter side then renders it with real components: timeline cards, checkable items, progress states.

This gets you three things:

  • Designable UI. A plan rendered as cards can be edited, checked off, and re-ordered. A plan described in a paragraph can only be read.
  • Validation. If the response doesn't parse into the schema, you know before the user sees anything. Retry silently, fall back, or degrade gracefully. A malformed chat message, by contrast, ships straight to the user's eyes.
  • Testability. You can unit-test everything downstream of the model with fixture data. More on that in Testing Flutter apps that talk to LLMs.

The chat surface still exists in Tertio — there's a coach you can talk to — but chat is the secondary surface. The primary surface is structured output rendered natively.

Context is the product

Two apps calling the same model produce wildly different value, and the difference is entirely in what you send with the request.

A bolted-on chatbot sends the user's message and maybe a system prompt. A real AI feature sends state: in Tertio's case, the user's stated goals, recent check-ins, what yesterday's plan looked like and what actually got completed, time of day, and the constraints they've set. The model isn't answering a question — it's operating on a live picture of the user.

Practically, this means the engineering work of an AI feature is mostly not prompt writing. It's:

  1. Deciding which state matters and modelling it properly.
  2. Building the pipeline that assembles that state into the request, every time, reliably.
  3. Keeping it small — context windows are big now, but latency and cost grow with every token, and irrelevant context measurably degrades output quality.

I'd estimate the prompt itself was under ten percent of the AI work on Tertio. The context pipeline was the feature.

Design for the failure modes

LLMs fail in ways traditional backends don't, and your UX has to absorb that:

  • Latency. A plan generation can take several seconds. Never block the UI on it — stream where possible, and where you can't, show meaningful progress against skeleton UI, not a spinner. Users will wait for something that visibly builds.
  • Bad output. Schema validation catches malformed responses; retry once before surfacing anything. For plausible but wrong output, give users a one-tap regenerate and lightweight edit controls. Editing an AI plan is not a failure state — it's engagement.
  • No connection. Cache the last good output locally. An AI planner that shows a blank screen offline is worse than a paper notebook.

The principle underneath all three: the model is a fallible remote dependency, and the app must stay useful when it misbehaves.

Keep the model swappable

I never hardcode a provider into the feature logic. The app talks to a thin service interface — generatePlan(context), coachReply(context, message) — and the concrete LLM lives behind it, usually proxied through my own backend rather than called directly from the device. That gives you:

  • key security (no API keys shipped in the binary),
  • the ability to switch or upgrade models without an app release,
  • one place to log, rate-limit, and measure.

Models improve monthly. Products that hardwired themselves to one specific model's quirks in 2024 are paying for it now.

What this looks like in practice

The stack that's worked for me across AI products, in one paragraph: Flutter front end with Riverpod for state; a backend endpoint that owns the prompt, assembles context, calls the LLM, validates the response against a schema, and returns typed JSON; the client renders that JSON with real components and treats regeneration as a first-class action. Chat exists where conversation is genuinely the right interface — coaching, clarifying — and nowhere else.

None of this is exotic. It's the same discipline you'd apply to any unreliable external service, plus one product-level rule that changes everything downstream:

Don't add AI to your app. Find the job the AI should own, and build the app around that.

Tasaddaq Hussain
Flutter developer & AI app engineer
Work with me
Next article
Generating a runnable route of exactly 42 km→