Field notes from three AI builders
What we learned about product decisions from testing Factory, Kiro and Atoms
We tested Factory, Kiro and Atoms on the same diary-product idea. The results showed meaningful planning, different delivery outcomes, and a sharper question about how owners understand the product they approve.
Orgenix field note · September 2026
In September 2026, we used Factory, Kiro and Atoms to explore a diary application that connects people through related thoughts and interests. We reviewed their questions and planning artifacts, followed implementation and independently checked the resulting applications where we could access them.
We did this while developing Orgenix, so we have a product thesis of our own. That makes it particularly important to give the other tools credit for what they demonstrated and to distinguish our interpretation from the observed results.
Two experiments with different starting points
The first experiment began with a relatively developed direction brief. It already described a private diary, system-selected matches, connections and explicit sharing. The evaluator answered follow-up questions, selected or approved recommendations and supplied constraints. Kiro and Atoms later received a change to the connection rules.
That experiment tested delivery and change with substantial product intent already supplied. It could not establish how well the tools would discover that intent from a vague idea.
The second experiment began with a much shorter idea: a diary connecting people through related thoughts and interests, using text, images and video. Each tool received the same opening request for a private web proof of concept using synthetic content, without purchases or public deployment.
The subsequent interactions were not identical. The evaluator asked a common follow-up for a concrete user journey and unresolved decisions, accepted different recommendations and corrected specific departures. Those interventions shaped the results.
These were small exploratory trials, with one run per tested workflow. They were not a controlled benchmark, a novice-user study or a ranking of current product capabilities.
Planning was substantive in all three tools
Factory's less specified run produced architecture, behavioral contracts, feature sequencing and acceptance assertions. It raised questions about visibility, consent, matching participation and sharing. After a follow-up, it described a concrete journey that made its proposed policies easier to discuss. The run was closed before an application was delivered.
Kiro presented requirements, revised them after clarification, then produced design and task artifacts. Its design explained how captions could supply media signals and how content visibility would change after connection. The evaluator approved that model.
Atoms used structured questions, a reviewable plan and an approval step. It recommended a multi-persona prototype and explained why it would demonstrate connection better than a single-user version. The evaluator later corrected an out-of-scope admin feature, and the observed app retained that correction.
The useful distinction is whether the right question arrives with enough context for the owner to answer it, and whether the answer remains authoritative as work proceeds.
Delivery results depended on the experiment
The two runs must be kept separate. Successful independent testing from an earlier application cannot verify a later application that we could not run.
| Tool | More specified experiment | Less specified experiment |
|---|---|---|
| Factory | Private diary foundation checked; matching, connections and sharing absent at the budget checkpoint | Substantial planning artifacts; no application delivered before the run closed |
| Kiro | Diary, matching and sharing ran independently; evolved connection rule and migration later checked | Completed build and tests reported; incomplete source transfer prevented independent execution |
| Atoms | Core diary and sharing flow plus a later change exercised with fresh accounts | Accessible preview exercised with prepared media fixtures, matching updates, reactions and persistence |
The outcomes came with different execution environments, scope decisions and assistance. We do not infer a fair speed or total-cost ranking from them. Prepared media fixtures did not establish arbitrary media understanding, and a vendor completion report remained unverified where the application could not be run independently.
A working change can leave an inconsistent product record
The earlier experiment asked Kiro and Atoms to replace request-and-accept connections with independent mutual interest. Existing connections and explicit shares were to survive; one person's old pending request must not become consent from both people. This was an evaluator-supplied change, so it does not show either tool discovering the preference spontaneously.
Kiro initially treated unanswered defaults as approved and began with an unwanted pending-request policy. The evaluator stopped and corrected it. The final exported implementation passed independent migration and behavior checks, but its README still instructed the operator to accept a request.
Atoms' evolved flow worked in independent fresh-account checks, while the credit-limited checkpoint left documentation cleanup unfinished. Controlled verification of some legacy migration behavior was incomplete.
Implementation correctness and consistency of the surrounding product record need separate checks. Linked documents do not guarantee that all of them describe the current system.
Approval does not tell us how much an owner understood
A document can accurately describe who sees a diary entry. A better decision experience may involve showing what a recipient sees, explaining the tradeoff and asking the owner to choose at that moment.
Our trials showed pieces of this experience: a concrete proposed journey, consequential questions, sequential artifact review and a working preview. They did not establish a sustained loop of digestible artifacts helping an ordinary owner discover implications and revise intent throughout delivery.
The decision-experience rubric was applied retrospectively, and we did not measure owner comprehension or reading time. Nor does more questioning necessarily improve the experience. Asking an owner to approve routine engineering details can add work without improving the product decision.
When evaluating an AI builder, ask it to walk through one consequential choice: what does the behavior mean for a user, what does it recommend, what changes with the answer, and what evidence shows that answer reached the result? Then inspect the result. A clear explanation and a functioning application are both necessary.
What this means for Orgenix
These experiments sharpened the experience we want to build: help owners understand an emerging product, decide in context and keep those decisions connected to subsequent work.
Orgenix is unfinished, and we are currently optimizing its architecture stage. We have not demonstrated superiority, lower customer effort or better end-to-end delivery through these trials. Those outcomes require their own evidence.
The immediate design question is specific enough to work on: how can the product make a consequential choice understandable without asking its owner to become the business analyst, architect and test engineer for every feature?
We develop that question in Who is running the AI engineering team?
Method note: these findings describe recorded September 6–7, 2026 checkpoints, with synthetic data, evaluator assistance and differing approved configurations. They are not a comprehensive security or accessibility audit. We did not measure comparative owner comprehension, total cost or time savings. Where an application could not be independently executed, its reported completion remains unverified.
First cohort opens in November
Follow the build, one useful update at a time.
Join the list and tell us how you are building with coding agents.