Systems & Infrastructure

AI Harness

A native iOS app that talks to a server I built and run myself, which routes messages to a language model, streams the answer back word by word, and now remembers what was already said. The interesting part isn't the app, it's everything that happens in between.

The foundation for a system I'm still actively building on, not a finished chatbot.

Chapter 01 · llmPrompt → Chapter 02 · llmChat
Role Solo, Client & Backend
What it is Native app + self-hosted backend + AI routing layer + conversation memory
Stack Swift · SwiftUI · Python · Starlette · SQLite · AWS
Status Live, evolving

Where the message actually goes

On the surface this is simple: type a message on an iPhone, get a reply from an AI model. But I didn't want the app talking directly to an AI company's servers. Instead, the message travels to a server I own and run, which decides which model should answer, sends the request along, and streams the response back into the app as it's being written, not all at once at the end.

That middle piece, my own server sitting between the app and the AI, is the actual project. I'm calling it my AI Harness because it's infrastructure I plan to keep extending: more capable routing, tools, memory, and eventually more complex mobile AI workflows all get built on top of this same foundation.

3 layers a message passes through: the app, my server, the AI model
1 backend, built to work with any number of AI providers behind it
Live runs continuously on a real cloud server, not just my laptop

Following one message all the way through

Click through each step below to see where the message is at that moment and why that step exists.

YouiPhone, typing
→
iOS AppSwift + SwiftUI
→
My Serverrunning on AWS
→
AI Modelpicked by the server
→
Answerstreamed back, word by word

Why not just talk to the AI directly?

The easy version of this project would have the iPhone app talk straight to an AI provider. I didn't build it that way, on purpose. A phone app is something anyone can pull apart, so it should never hold private keys or make permanent decisions about which AI model to use. Instead, all of that lives on my server, in one place I control.

Keys stay off the phone

The credentials needed to call an AI provider live only on my server. The app never sees them, so there's nothing sensitive to extract even if someone took the app apart.

One place to change everything

Which AI model answers, what happens if one fails, how a future feature like memory or tools would work, all of that can change on the server without touching or resubmitting the iPhone app.

The app is the interface. The server is the control layer underneath it.

What's running on the other end

My server is a small Python program. I built it to be asynchronous, meaning it doesn't sit and wait around while an AI model is thinking. It can start one request, immediately move on to handle something else, and come back when the model has something to say. That matters because AI responses can take a few seconds, and a server that just freezes during that time would be a bad building block for anything more advanced later.

Three pieces work together here, each doing one job:

My CodePython, written with a framework called Starlette — defines what each request should do
↓
Granianthe actual web server — receives traffic from the internet and hands it to my code
↓
The InternetHTTP/2, a faster, more efficient way of moving data than older web traffic

In plain terms: Starlette decides what my server does with a request. Granian is the part that's actually listening for that request in the first place. Splitting those two things apart is a standard pattern in real backend systems, and it's the first time I built something that way instead of running one script that did everything itself.

Why the reply types itself out

An AI model doesn't produce its whole answer instantly, it generates it piece by piece. My server doesn't wait for the whole thing to finish before sending anything back. As soon as the model produces a small chunk of text, the server forwards that chunk straight to the app, which is why the reply appears to type itself onto the screen instead of showing up all at once.

Under the hood, this uses a technique called Server-Sent Events: my server keeps the connection to the app open and pushes small pieces of text down it one at a time, instead of closing the connection after a single reply the way a normal web request would.

Without streaming

Send the message, then wait, with nothing on screen, until the entire answer has been generated somewhere else and finally arrives all at once.

With streaming

Words start appearing within a second. The full answer takes the same amount of time to finish generating, but it never feels like the app is stuck doing nothing.

The parsing turned out trickier than the concept. SSE events are separated by blank lines, but Swift's built-in line reader silently skips empty lines, which means it throws away exactly the delimiter the protocol depends on. I ended up reading the stream byte by byte and building lines myself so I could actually see where one event ended and the next began.

Rein.swift, manual SSE line buffering
// AsyncSequence.lines skips empty lines, which SSE relies on
// as the delimiter between events, so lines are built by hand instead
var buffer = Data(capacity: 8192)
for try await byte in bytes {
    if byte == UInt8(ascii: "\r") { continue }
    if byte != UInt8(ascii: "\n") { buffer.append(byte); continue }

    let line = String(decoding: buffer, as: UTF8.self)
    buffer.removeAll(keepingCapacity: true)
    // an empty line here is the real "\n\n" event boundary

The app doesn't need to know who answered

My server isn't locked to one single AI provider. It keeps a list of providers it's allowed to use, starting with Groq, and picks one for each request. If that provider is slow, unavailable, or returns an error, the server can try the next one on the list instead of just failing outright. The app never has to know or care which one actually generated the answer.

iPhone Requestjust says "please answer this"
→
My Serverpicks the best available provider
→
Groq or OpenRouterwhichever actually responds

This stopped being theoretical fast. I started with Groq alone, and once I was actually testing regularly, I ran straight into its rate limit. Instead of just waiting it out, I added OpenRouter as a second provider the server could fall back to. That one annoyance is basically why the fallback logic exists in its current form, not because I planned for it up front, but because I hit the exact failure it's meant to handle.

Failure is part of the design

If Groq is rate-limited or down, the server falls back to OpenRouter rather than assuming the first provider always works. Only if nothing responds does it tell the app so honestly, instead of pretending it succeeded.

Swappable by design

Because the app only ever talks to my server, not to a specific AI company, I was able to add OpenRouter as a second provider without changing a single line of the iPhone app.

Setting up HTTPS taught me more than I expected

My server has its own address on the internet and only accepts encrypted connections, so no one sitting in between can read what's being sent. While building and testing this, I had to set up that encryption myself, which meant the app and the server had to be introduced to each other and agree to trust one another before any conversation could happen.

That taught me something I didn't expect going in: "just turn on HTTPS" isn't a single switch, it's an actual trust relationship between a device, a certificate, and a server, and every piece of that chain has to be set up correctly for it to work at all.

Turning a script into an actual service

I didn't want to have to manually start this server every time I wanted to use the app. So instead of running it as a script in a terminal window, I set it up as a proper background service on the cloud machine it lives on. That means it starts automatically whenever the machine boots, and if it ever crashes, it restarts itself instead of just staying down.

Server Bootsthe cloud machine turns on
↓
Service Managerautomatically starts my backend
↓
Running Continuouslyrestarts itself automatically if it ever crashes

That's a small change on paper, but it's the difference between "a script I run when I remember to" and "a service that's actually available." It's the same expectation any real product would need to meet.

Chapter 02 · llmChat

An LLM doesn't remember you

Everything up to this point could send a prompt and stream back an answer, but every request was independent. The model had no idea what had been asked five seconds earlier. If I wanted the app to feel like an actual conversation instead of a series of one-off questions, I had to solve a different problem: how do you build context around a model that has none of its own?

Without context

"Where is Tokyo?"sent alone
→
"Tokyo is in Japan."correct, but isolated
"What did I just ask?"sent alone, right after
→
No ideathe model never saw the first message

With context

Full historyboth turns, sent together
→
"You asked where Tokyo is."answered correctly

Language models don't automatically know what happened earlier in a conversation. Every time I wanted a coherent back-and-forth, I had to explicitly send the prior turns back to the model along with the new message. The question stopped being "how do I call an LLM" and became "how do I reconstruct a conversation for a model that forgets it the instant it answers."

I added memory without putting memory in the app

The app still just sends a message. All of the new work happens on the server: it identifies which app is asking, looks up what that app's conversation has said so far, hands the model the full history instead of a single line, and saves the new exchange once it's done. The app never has to manage any of this itself.

SwiftUI Appsends the new message + an app identifier
↓
harnessdloads prior turns, builds the full context, calls the model, streams the reply
↓
SQLitestores every turn so the next request can rebuild the conversation

Watch what actually gets sent to the model as a conversation grows. Each turn you add here gets appended to the request payload, the same way it happens on the server.

Conversation
What the server sends the model
{
  "messages": []
}

Nothing sent yet. Click the button to start the conversation.

The model never "remembered" the first message on its own. The trick is that the full history gets sent back to it every single time, and that history has to live somewhere between requests since the model itself keeps none of it.

harnessd, saving a turn and rebuilding history
# save the incoming prompt, tagged with appID, before calling the model
async with rwdb.cursor() as cur:
    await cur.execute(
        "INSERT INTO turns (appID, prompt) VALUES (?, ?)",
        (openAIRequest.appID, openAIRequest.messages),
    )
    turnID = cur.lastrowid

# then pull every turn that belongs to this appID, oldest first,
# and hand the model the whole thing as its message history
async with rwdb.execute(
    "SELECT turnID, prompt, reasoning, completion FROM turns "
    "WHERE appID = ? ORDER BY turnID ASC", (appID,)
) as cur:
    turns = [Turn(*row) for row in await cur.fetchall()]

One backend, more than one conversation

My server isn't only ever going to talk to one app. That means it can't just pull "the most recent conversation" out of the database. Every stored turn is tagged with an app identifier (the iOS app's bundle ID), and the server only loads turns that match the app making the current request before it builds the context to send the model.

appIDedu.umich.mahimash.Agent
→
Filter the tableonly turns tagged with this appID
→
Just this conversationsent to the model as context

It's a small field doing an important job. Without it, one shared table of stored turns would turn into one shared, tangled conversation. With it, the same backend and the same database can hold as many separate conversations as there are apps talking to it, and each one only ever sees its own history.

Two models, the same memory layer

These are real screenshots from the simulator, not mockups. Both runs ask the same follow-up question, "what did I just say," after a first message, using two completely different models behind the same backend.

iOS simulator showing the llmPrompt app before any message has been sent

Before: no conversation yet

iOS simulator conversation with gemma4:e4b correctly recalling the first message

After: "You just said, 'howdy?'"

iOS simulator conversation with qwen3.8-27b replying to the first message

First message, answered by qwen3.8-27b

iOS simulator conversation with qwen3.8-27b after the follow-up question

Follow-up sent to the same conversation

Because the app never talks to a provider directly, swapping which model answered a given conversation doesn't change anything about how that conversation's memory is stored or retrieved. The context layer sits below the model, not inside it.

Every part is answering a specific question

Why SQLite?

Because conversation history has to live somewhere between one request and the next. The model doesn't hold onto it, so something persistent has to.

Why an appID?

Because the backend is shared. Without a way to separate conversations, one app's history could bleed into another's.

Why SSE?

Because a model generates its answer piece by piece. Streaming lets the app show those pieces as they're produced instead of waiting on all of them.

Why a backend at all?

Because provider credentials, routing decisions, and now conversation storage shouldn't live inside a mobile app that anyone can take apart.

Why async?

Because the server shouldn't sit idle waiting on one model to finish generating while a second request is trying to come in.

Why systemd?

Because a server that only runs while a terminal window happens to be open isn't a server the app can actually depend on.

llmPrompt proved I could connect an app to an LLM. llmChat turned that connection into a system that could hold an actual conversation.

What actually stuck with me

The interface and the control layer are different jobs

The app's whole job is showing a conversation and reacting to typing. Every decision about credentials, which model answers, and what happens if something fails belongs on the server. Keeping that boundary clean made both halves easier to reason about.

Streaming changes how you have to think about a response

A reply isn't one thing that arrives, it's a sequence of small pieces arriving over time. The app has to be built to update itself as new pieces show up, not to wait for a single finished answer.

Networking is layers, and each layer solves a different problem

Encryption, the actual web server, my application code, and the way requests get routed all solve separate problems that stack on top of each other. Understanding where one layer's job ends and the next one's begins made debugging far less mysterious.

Getting code to run on my laptop and keeping a service reliably available on a machine I don't touch every day turned out to be two very different skills.

Where this goes from here

Right now the system can securely move a message from a native app, through my own server, to an AI model, and stream the answer back. That's deliberately the simplest possible version of a much bigger idea. I'm continuing to build on this same backend as I explore giving it tools to use, longer memory of a conversation, and more capable mobile AI experiences.

Current status

The system can securely move a message from a native app to my own server to a model and back, stream that answer in as it's generated, and remember the conversation it's part of. Everything below is what I'm actively building toward next on top of the same foundation.

Tools the model can use Memory beyond one conversation Multi-step agent workflows Images & other input types More capable mobile experiences

The tools behind it

The system spans a native SwiftUI client, a Python backend, SQLite persistence, model routing, and SSE streaming. I wrote the client, the server, the conversation storage, and the routing logic myself, with the AI providers handling only the actual language generation.

Swift SwiftUI Python Starlette Granian HTTP/2 Server-Sent Events AWS systemd

Next project

A2B

View case study