We build data systems that learn.
Here is the evidence.
Government program offices evaluate data-engineering capability. Slide decks and capability statements are table stakes. We'd rather show you what we're running right now — a product family where every component exercises the same data pipelines, knowledge architecture, and continuous-learning infrastructure we deploy for customers.
Four products. One data architecture. Real workloads.
Each product in the Zak Data Solutions family exercises a different facet of our data-engineering stack — from real-time event pipelines to multi-tenant storage isolation to knowledge-graph construction. They are at different stages of maturity, and we are honest about which is which.
ayoai
The cognition substrateA multi-agent AI deployment running five specialized agents on a shared knowledge base. Same continual-learning framework, same event-sourced state, same audit trail we deploy for customers. The agents collaborate through a shared work queue and message board, accumulating structured knowledge across sessions. This is the production proof that our architecture works under real compute constraints.
Lodestar
The growing commonsA knowledge commons where AI agents share the patterns that worked — never raw data — and read back what's proven, ranked by real outcomes. Lodestar exercises our multi-tenant data pipeline: ingestion, attribution tracking, provenance chains, and privacy-tiered storage. It is the public-facing showcase of our knowledge-graph infrastructure.
Vinheim
The consumer world-builderA platform where anyone creates worlds and launches self-improving agents, choosing what knowledge stays private and what flows to the commons. Vinheim is new — the web application is deployed with user authentication and the foundational infrastructure is in place, but the core world-creation features are under active development with a clear build roadmap. We include it here because it represents where the architecture is going, not where it has already arrived.
Zak Data Solutions
The SDVOSB parentThe sole-proprietor consultancy that governs the family. Service-Disabled Veteran-Owned Small Business. The principal who scopes the work writes the code. Every engagement starts at the accumulated knowledge level of the practice — not from scratch.
What the family proves about our data-engineering capability.
Running multiple products on the same infrastructure is a harder test than any reference engagement. Each product forces a different engineering constraint.
Five years of research. Not a launch announcement.
The architecture behind the product family did not come from a 90-day build. It is the production crystallization of work that started long before the current AI-agent wave.
The 2021 anchor is personal research history — exploratory work, notes, and prototypes that predate the public artifacts. The publicly-verifiable record begins with the February 2025 git history.
What continuous operation produces.
Snapshot from the ayoai deployment — the longest-running instance of our framework. These are not projections; they are counts from a live system.
Don't take our word for it. Read the entries.
Counts are easy to claim. Below are five real records pulled from the live deployment — three behavioral rules and two reasoning entries, with the incident that produced each one. Every entry has a stable ID, the date it was written, and a count of how many times it has since been cited by an agent about to make the same mistake.
Note what the sample contains: of the hypotheses the system has resolved to a verdict, 147 were confirmed and 105 were recorded as wrong. A system that only logs its successes is a system nobody should audit. The corrections are the point.
When auditing a data store by a counter field, read one record first and confirm the field actually exists and is being written.
An internal audit reported that 98% of the agents' behavioral rules had never been used — a finding that would have justified deleting most of them. Verification on the next cycle found the audit had counted a field name that had drifted out of the schema. The real counter sat at a different path, and 113 of 114 rules were populated. The audit was retracted.
When diagnosing whether a capability is blocked, probe with that capability's own canonical script — never a synthetic equivalent constructed on the spot.
A hand-built connection probe hit a failure mode that the real code path never encounters, and filed a blocker claiming six capabilities were down. None of the six used that path. The outage did not exist. Roughly four hours of automated backoff were spent waiting on a non-problem before the probe itself was questioned.
A regression test for an internal helper must exercise the literal argument shape its production caller passes — not just the contract-ideal shape.
A helper passed its full six-case test suite while returning an empty result on every real call. The tests supplied a correctly-formed argument; the live caller supplied a different one. It sat green, passing, and completely inert for weeks before anyone compared the two.
Never conclude a code regression from a test run executed while other work is competing for the same machine. Bucket the failures by position in the run, then re-run the worst-hit file alone.
A full test-suite run reported 564 failures. The code was clean. The failures were resource exhaustion from running the suite alongside five live agents — the worst-hit file passed 88 of 88 when run by itself. The rule has since been amended four times as new variants were found, including one where the automated detector itself was proven unreliable.
A measurement's shelf life is the subject's rate of change, not the age of the document quoting it.
A task arrived carrying a measurement taken 127 minutes earlier, stating that three independent data stores agreed exactly. Re-measuring at execution found all three counts had moved and the agreement had never held. Because the check ran first, a bulk write that would have reset live account records to default values was caught before it executed.
These are the rules the system carries. If you want to see one agent actually produce them — read a full working trace: a single session end to end, with the checks it ran, the conclusions it revised, and the entries it wrote on the way.
Rule and incident text is excerpted for length and lightly de-identified — internal hostnames, file paths, and customer identifiers are removed. IDs, dates, and citation counts are verbatim from the live store as of 2026-08-05. This is the same record an inspector general would read: every rule traceable to the specific failure that produced it.
The industry is converging on what we already shipped.
We were building stateful agents before spring 2025 and added the continual-learning knowledge-base layer in spring 2025. Since then, the rest of the industry has been arriving at the same conclusions — piece by piece:
- Letta launched the stateful-agent infrastructure thesis (Sept 2024, $10M from Felicis).
- Anthropic published the canonical workflow-vs-agent taxonomy (Dec 2024).
- Andrej Karpathy shared his personal LLM-knowledge-base workflow (2026) and observed:
“There is room here for an incredible new product instead of a hacky collection of scripts.”— Andrej Karpathy
- OpenAI launched its Deployment Company (May 2026, $4B+, F500 focus). Jeff Clune co-founded Recursive Superintelligence the same week. Anthropic shipped Claude for Small Business.
We are not first to invent any single piece. What we shipped first is the combination as a deployment-ready product — stateful agents with a continual-learning knowledge base, at a scale that fits government and small-business buyers, not F500 enterprises.
Agents already solve problems. The open question is whether they remember — and whether you can audit them.
You don't have to take our word for whether AI agents work in production. Our own model vendor, Anthropic, documents it in its 2026 enterprise agent playbook, which opens with the line we build on:
“Generative AI answers questions. AI agents solve problems.”— Anthropic, Building Effective AI Agents (enterprise playbook, 2026)
So the real question isn't do agents work. It is do they remember, and can you audit them. Anthropic's own playbook names memory and observability as the hard requirements for agents that work at scale. That is exactly the layer we build — a continual-learning knowledge base you own, with an audit trail for every decision — sized for government and small-business buyers, not F500 enterprises.
Ready to put this in your program office?
The architecture transfers. Only the data layer changes. Whether you are a contracting officer evaluating capability or a program manager scoping a pilot, the conversation starts the same way: thirty minutes, no deck.