Work · Platform ยท Self-hosted AI

Nick AI platform

The AI system behind this site: long-context processing, vector memory, a site guide, an editorial pipeline and a live map of what it knows, with judgment and security layers in front of every model call.

Role Founder, architect and builderWhen 2025โ€“2026
Nick AI platform
~57,000lines of Python in one FastAPI app: the site, the guide, the pipeline and the security layers
10,000 charsthe point where a request switches to RLM, which works through long context in pieces
Hourlyrebuild of the live knowledge map: clustered, named and laid out in 3D
4 levelsin the gatekeeper: normal, slowed, paused for an hour, blocked for a day

The problem

Most AI on a personal website is a chatbot bolted onto a page: a model with a system prompt, no real knowledge of what the site says, and nothing between a stranger's text and the model. It can't cite the site's own work, it forgets everything, and it will happily be talked into anything.

I wanted the opposite: an assistant that knows the site's own material, a pipeline that helps me research and write, and security designed the way I'd design it for a client.

The idea

Nick AI is one platform that runs this site. It keeps the site's knowledge (articles, research sources and documents) in vector memory. It answers visitors through the Nick AI chat and the site guide, the orb on every page. It runs an editorial pipeline that researches and drafts, but publishes only what has been approved. And it shows what it knows as a live map on the home page.

Models are services the platform calls, not the platform itself, so every call has a judgment layer and guard rails in front of it.

How it's built

It is a single FastAPI application serving the site's pages, about 57,000 lines of Python, running in Docker behind Cloudflare. Knowledge lives in Qdrant, embedded with all-MiniLM-L6-v2 into separate collections for articles, research sources and documents, and an ingest loop refreshes it every hour. Operational state lives in SQLite.

Answers come from Llama 3.3 70B, run by a private, no-log provider that doesn't store prompts or answers, with a local Ollama model as the fallback. In front of the main model sits Jev from TypeSafe, a fast judgment layer that triages what a message is actually asking for before anything expensive happens. Kokoro provides the voice for spoken articles.

The live neural map on the home page: the platform's memory clustered into topics, rebuilt every hour.

RLM: working past the context limit

Long inputs are handled with a Recursive Language Model workflow, based on the RLM research paper. Once the context passes 10,000 characters, the model doesn't try to read it all at once. It writes small Python programs in a restricted sandbox to inspect the material, and can call itself on the parts that matter, for up to ten rounds.

If that doesn't converge, the platform falls back to plain chunking: 50,000-character pieces with an overlap, each processed on its own, then combined into one synthesis. The point is to get an answer grounded in the whole document, not just the part that fitted in the window.

How Nick AI works: RLM for long context, Qdrant for memory, and the workflows around them.

The live brain

Every hour a job reads the public knowledge collections, clusters them (k-means, with the number of topics chosen by silhouette score), lays them out in 3D and has the model name each topic. The home page renders the result as a live neural map you can rotate and play back over time.

A second layer shows themes from the X account's activity, as themes only, never individual posts. Snapshots are kept for 48 hours, and if the builder is down the site keeps serving the last good one.

The site guide and its gatekeeper

Before a visitor's message reaches a model it passes a series of cheap checks: a hidden honeypot field, bot user agents, messages typed impossibly fast, repeats, and rate limits of 8 messages per 5 minutes and 40 a day per visitor. Then Jev screens the intent, and PromptGuard checks for injection patterns going in and redacts keys, database URLs, internal paths and private IP addresses coming out.

A gatekeeper scores every signal, with a 30-minute half-life so a single bad moment fades. Enough of them slows a visitor down, then pauses the guide for an hour; only scanner behaviour can block someone, and then for 24 hours. Guide conversations are kept in memory for context and expire after three hours idle, the guide only ever links to pages on this site, and a daily retention job deletes stored data on the schedule the privacy policy sets out.

The editorial pipeline

The platform keeps a small queue of article ideas, researches them daily and drafts with the model, keeping the source evidence alongside each draft. Nothing publishes itself: only drafts I have approved go out. The X account is run by a separate automation stack; the site only reads its activity to build the map.

Where it doesn't fit

Self-hosted here means the platform, the memory and the data. Most answers come from a hosted open model (Llama 3.3 70B on a no-log provider), with a local model only as a fallback, so this is not a fully local AI. The Jev intent screen is designed to fail open: if the judgment service is unreachable the site keeps working, and the rate limits, PromptGuard and the gatekeeper still apply, but that screen doesn't.

Delegating work to external agents is the part still being wired back in, so the chat answers from the platform's own engine today. The knowledge map is statistical: the clusters overlap and the topic names are written by a model, so it's a picture of what the memory holds, not a taxonomy. And like any AI, the answers can be wrong, which is why the guide points to the pages themselves.

What I took from it

The model is the least interesting part. What makes it trustworthy is everything around it: memory the answers can be traced to, cheap checks before expensive ones, a judgment layer before the model, redaction after it, and a person approving anything that gets published. It's the same fail-closed thinking I bring to identity and AI governance work, applied to my own site.

Building something that has to be secure by design?
I help teams get identity, secrets and AI access right from day one.
How I work with teams →