Falcra Instant by Falcra

Instant by Falcra

Instant Practices

How Instant runs the AI-Driven Development Lifecycle safely: the decisions, foundations, assurance and AI rules behind it, in detail. It is the reference for the crew, your architects and your technical owner.

The Instant Guide describes what Instant is. And the practices below describe how its parts are carried out. All of them happen alongside the build, in the room, over a few days.

To follow them step by step: the Instant run sheet puts every step in order, with who leads it and when it's done.

Published by Falcra.
The Instant Framework (theinstantframework.com) by Falcra (falcratechnologies.com)

The Instant Framework is published in three parts: the Manifesto (values and principles), the Guide (what Instant is, its roles, stages and gates) and the Practices (how its decisions, foundations, assurance and AI rules work in detail).

Chapter 1

Choices you make

In short

Every engagement involves a handful of decisions that belong to your organisation, not to the crew or the AI. Each is put to the room, made by the person named below, and recorded.

Minutes, not meetings. These choices are made in the room during preparation and the Blitz, by the people who own them, with the crew bringing the options ready to decide.

ChoiceOptionsWho decidesWhen
CrewA Combined, B Split, C Developer ConductorYour business ownerS0
Blitz formatConsecutive days, or split into sessions (Guide, chapter 6)Your business ownerS0
Requirements or user storiesOne of the two, depending on your organisationYour business owner, advised by the ChallengerS2
Data architecture level1 Standard, 2 Standard plus audit trail, 3 Event SourcingYour business owner, after six questions put to the room. The crew advises but never chooses.S2
Risk tier for each part of the applicationCritical, Important, StandardYour experts and business ownerS2
Security constraints packageOptional; written by your security leadYour security leadS0
JiraOptionalYour business owner or project managerS0

Requirements or user stories

The AI build tool offers two ways of capturing what the software must do: requirements, or user stories. Generally only one is needed, so only one is used. Which one suits you depends on how your organisation already works. Either way, the business reviews the result in the room (sign-off G1).

Risk tiers

Not every part of an application matters equally. A mistake on a help screen is an inconvenience. A mistake in a payment calculation can be a loss, a fine or a safety event. Instant asks the room to sort every part of the application into one of three tiers, and to record who decided and when.

Critical

Safety, regulatory, financial or revenue logic. Anything where an error causes serious harm.

Important

Core business workflows and data integrity.

Standard

Screens, layout, forms and plumbing.

Recording and managing tiers

Tiers are kept in the risk tier register: one row for each part of the application, with its tier, the reason, who decided and when. The AI build tool proposes the parts and a suggested tier for each, so the room starts from a list, not a blank page. Your experts and business owner confirm or change each one, the Challenger tests the reasons, and the Conductor keeps the register. No row may be left undecided at gate G2. When scope changes, rows are added or re-tiered, and each change is logged beside the original decision. The register goes into the evidence pack.

A ready-made register (a spreadsheet with the tiers, reasons, change log and a summary) is a free download from theinstantframework.com.

Why it matters. In most applications only a few parts are Critical: a handful of rules and calculations. Line-by-line human review is concentrated there, while automated checks and AI review cover the rest (chapter 3). That is how review keeps pace with AI-written code without lowering the bar where it counts.

Data architecture: how much change history to keep

This choice is about the record of changes to your data: who changed what, when, and what it was before. It is not about the business data itself, which is stored in full at every level. If the application holds transactions, cases or readings, those are kept whichever level you choose.

The level is chosen in Specify, before design, because it shapes the database design and is costly to change later. Instant's data architecture rule tells the AI build tool to stop before design until a level is recorded, and gate G2 can't be passed without it. The crew advises; your business owner decides.

An example: a manager raises an approval limit from $5,000 to $10,000.

Level 1

Standard

The record holds its current value: $10,000. The data itself is stored and maintained as normal; what isn't kept is the change, so the earlier value and who changed it are not recorded. Suits internal tools, simple data entry, reporting and proofs of concept.

Level 2

Standard plus audit trail

As Level 1, plus a permanent audit trail of every change: the limit went from $5,000 to $10,000, changed by the manager, with their role at the time, at a recorded date and time. Suits most enterprise business applications.

Level 3

Event Sourcing

Every change is stored as an event, in order, and never altered, and the current record is worked out from its events: the limit is $10,000 because it was set at $5,000 and later raised. Any record can be rebuilt as it stood on any past date, under the rules in force at the time. Suits regulated and audit-critical applications.

Level 3 has a cost: more design up front, a specialist pattern your own team must be able to maintain, and changes that take a moment to appear on screen. Nothing is deleted at Level 3, so where privacy law requires personal data to be removed, the design handles it explicitly.

Recording who viewed what, as opposed to who changed it, is access logging. It is a separate decision, often required for personal or sensitive data, and can be added at any level.

Six questions to choose the level

  1. Will a regulator, auditor or court ever ask why something was decided, or what was known at a past date?
  2. Do the rules that drive decisions change over time, and must old decisions stay explainable under the old rules?
  3. Will several people change the same records at the same time?
  4. Must the change history of a record be visible in the application itself?
  5. Who will maintain the application after handover, and have they worked with Event Sourcing before?
  6. Is there personal data that may have to be deleted on request?

A yes to question 1, 2 or 4 points toward Level 3. A no to all of questions 1 to 4 points to Level 1 or 2. Question 5 weighs against Level 3 if the people who will maintain it have no experience with it. The crew shows the room what the answers point to, and your business owner decides. The level, the answers, who decided and the date are recorded. If the level changes later, the change is recorded alongside the original decision, not instead of it.

The security constraints package

Some organisations have security rules the build must never break. If you want that, your security lead writes them down as a short set of plain files before the Blitz. Examples: the application must not be reachable from the internet, people must sign in through your identity provider, only approved protocols may be used.

The AI build tool is given the package and builds within it throughout. If someone in the room asks for something that would break a constraint, the build stops and says it can't proceed. It does not look for a way around. A halt isn't overridden in the room. Only your security lead can change the package.

The AI build tool has its own generic security defaults. They are standard industry practice, not your policy. The package is how your actual policy gets into the build.

Jira

If your organisation uses Jira, the AI build tool can create cards there during the engagement, through a Jira connector. The cards are created on the Stage Manager's instruction, or the Conductor's where there is no Stage Manager. The starting point is the Instant Jira backlog: 15 epics and 130 tasks covering everything from the first conversation to handover, including the approvals, security, environments and production readiness that usually stall a project. Each task carries its stage, an owner role, and whether it is always needed or only in some situations, and long-lead tasks are flagged to start during preparation. It imports straight into Jira. It is free under the Creative Commons Attribution 4.0 licence (CC BY 4.0): you can adapt it and share your version, as long as you credit Falcra. It is a free download from theinstantframework.com.

Information Model and Data Model

When the application is built, the AI produces an Information Model for your architects: what information the application holds, what each item means, where it comes from, who owns it, how it is protected, how long it is kept and how it maps to the database, with diagrams. Every definition says where it came from: an approved requirement, the code, the running application, or an assumption still to be confirmed. Nothing is filled in from what applications like this usually contain, and gaps are listed with the role that should answer them. The AI reads the requirements, decisions and code and changes nothing. If you supply your enterprise information model or glossary, it maps to that. The skill and template work with any application, are free under the Creative Commons Attribution 4.0 licence (CC BY 4.0), and are a free download from theinstantframework.com.

Alongside it, the AI produces a Data Model for your database administrators, developers and support teams: how the information is stored. It covers every table, column, key, constraint and index, any views, procedures and triggers, how the chosen level of change history is implemented, database security, and the migration history, with diagrams, a data dictionary for your catalogue tools, and a schema export. It is generated from a non-production copy of the database itself, ideally through a direct read-only connection (chapter 2), so it is accurate on the day it is produced, and it is regenerated at every release. Where you supply database design standards, it lists any departures; without them, it describes and doesn't judge.

The two are kept apart because they answer different questions for different people: what the information means and who owns it, for architects; and how it is built, for the people who run it. Produced together, they reconcile: every entity maps to its tables, and every business table to an entity. The Data Model skill and template are also a free download from theinstantframework.com.

Chapter 2

Enterprise foundations

In short

Four areas need settling in any enterprise build: non-functional requirements, existing data where there is any, sign-in and access, and environments. Your own standards come first in each. Instant adds a disciplined way of applying them at AI speed, and a starting point only where you have none.

Settled in the room. These are decided during the Blitz, mostly in Specify and Design, alongside the build. Only items with long lead times, such as environments and data access, start earlier, during preparation.

10.1 Non-functional requirements

Your standards come first. Most large organisations have a non-functional requirements catalogue, architecture standards, or service levels set by application criticality. These are given to the AI build tool in Specify and Design, and they drive its own detailed non-functional work. In AI-DLC, that is the NFR Requirements and NFR Design stages, run for each Unit of Work during Build.

What Instant adds is the business half of the conversation. Availability, recovery and acceptable data loss are business decisions, yet they are often left for technical teams to assume. In Specify, the Conductor puts ten questions to the business owners in the room, so they state those answers themselves, on the record. The answers feed your standards and the engine's NFR stages. They don't replace them.

#Question for the businessWhat it informs
1How much of the time must it be available, and when are people using it?Availability target and support hours
2If it goes down, how long can the business wait for it to come back?Recovery time objective
3How much recent data could the business afford to lose?Recovery point objective and backups
4How quickly must screens respond for people to stay productive?Performance targets
5How many people use it at once, normally and at peak, and how will that grow?Capacity and load testing
6Who may see or change what, and is any of the data sensitive or personal?Access control and privacy
7How long must data be kept, where may it be stored, and how is it deleted?Retention and data residency
8Who is told when something breaks, and who fixes it?Alerting and the support model
9Who needs to be able to use it, including people with disabilities?Accessibility standard
10What may it cost to run each month, and who pays?Cost budget and alerts

Where there is no standard. Smaller organisations, or a first application of its kind, may have no catalogue to draw on. For those, Instant offers a starting set scaled by risk tier. The targets below suit the Important tier. Critical parts get tighter targets agreed with your business owner, and Standard parts may relax them.

MeasureStarting target
Availability99% in a calendar month
Recovery time8 hours
Recovery point24 hours
Everyday actions95% within 0.5 seconds; 99% within 1.5 seconds
Rollback of a bad release15 minutes
Detection of a critical failure5 minutes
AccessibilityWCAG 2.2 level AA

Whatever its source, every target is proven by a test and signed off in the evidence pack, or recorded as an exception with an owner and an end date. Your security constraints package, if you have one, overrides any target wherever it is stricter.

10.2 If your project involves existing data

Many Instant builds are new applications with nothing to bring across. If yours is one, skip to 10.3. Where the new application replaces, or draws on, a system that already holds data, migration is a workstream in its own right. It follows your organisation's data and migration standards, runs alongside the build, and often continues well beyond the Blitz. Instant adds a few disciplines to it.

  1. 1

    Protect first

    A backup is confirmed in writing, and the crew works through a read-only account, ideally on a copy of the database. Who: your database administrator or system owner.

  2. 2

    Start from the new application

    Mapping is driven by what the new application needs: for each field, where its data comes from. Anything not needed is left behind deliberately, with the business's agreement. Who: the Conductor and the AI, with your experts deciding.

  3. 3

    Discover from evidence

    The AI works through the documentation, then the database's own metadata, then the data itself, confirming what each field holds before anything is built on it. Who: the Conductor, through the read-only account.

  4. 4

    Rehearse the migration

    The AI writes the migration and reconciliation scripts. They are rerun against a copy until clean, and timed so the cutover window is known. Who: the Conductor and Roadie, with your database administrator.

  5. 5

    Prove it

    Counts, totals and samples are reconciled, rejected records are decided by their owners, and the results go into the evidence pack. Who: your data owners and experts.

Letting the AI query the database directly

The fastest way to work with an existing database is to connect the AI build tool to it directly, through a read-only database connector: an MCP server, the standard way AI tools connect to other systems. The AI then runs its own queries, reads the results and moves on, in minutes rather than the hours or days it takes to send queries to a DBA and wait for the results. The same connection serves wherever the build needs to read existing data: discovery for a migration, checking reference data, a system the new application will read from, and producing the Data Model (chapter 1).

  • Your approval first, in writing, with the backup confirmed.
  • A read-only account. Your DBA creates a dedicated account that can read but not change anything, ideally on a copy of the database, and proves it with one harmless write that fails.
  • Credentials stay out of the AI's view. The connection is stored as a saved connection or in environment settings, never typed into the conversation or committed with the code.
  • Every query is visible. Each query is shown and approved before it runs, so your DBA can see exactly what was asked of the database.
  • Connected only while needed. The connector is added for this project alone and removed when the work is done.

Where your policies don't allow a direct connection, the AI writes each round of queries as one list, your DBA runs them, and the results come back together. It works, but it is much slower, so a direct read-only connection is worth requesting during preparation.

Working with the AI on existing data

  • Read-only is enforced by the account, not by instruction. Instructing the AI to only read is a second layer. The control is a database account that cannot write, proven by one harmless write that fails.
  • Work on a copy where possible, so discovery queries don't load the live system.
  • Profile before viewing rows. The AI sees whatever a query returns, so personal and sensitive data is profiled with counts and summaries, and individual rows are viewed only when needed.
  • Column names are clues, not facts. An AI reading a schema will reasonably infer meaning from names such as "status" or "active". Left unchecked, an inference like that can travel into the design as if it were confirmed. Instant has each one confirmed by a query, or by someone who knows the data, and marks it unverified until then (chapter 4).
  • Ask once, answer once. When queries go through a person rather than a direct connection, the AI asks for each round as one numbered list, and the results come back together.

The mapping

The mapping is kept as a spreadsheet with one row per field in the new application: its source, the transformation rule, and its status (verified, needs a decision, unverified, or new with no source). A separate list records what is left behind, why, and who agreed. The migration scripts and checks are generated from the mapping, so the two can't drift apart. Your experts own what the old data means and what is left behind.

Logic held in the database

Where the existing database carries business logic in stored procedures, triggers or views, that logic is a scoping question in its own right. In many organisations it is extensive, complex and undocumented. The AI can inventory and summarise it from the database itself, which makes its scale visible in Specify rather than in testing. Your experts and technical owner then decide what the new application must reproduce, and that work is planned explicitly.

Data quality

For each data quality problem found, the data owner chooses where it is fixed.

Fix itWhen it suits
In the old system, before migrationFew records, and the old system is still in use.
During migration, by ruleA pattern that can be corrected by rule, such as formats or codes.
In the new system, after go-liveIt needs judgement record by record. The rows are flagged for follow-up.

An unmade decision turns into rejected records on go-live day.

Reconciliation

Every rehearsal, and the final run, is reconciled:

CheckPasses when
CountsFor each kind of record, the number in the old system equals those migrated plus rejected plus deliberately left behind.
TotalsKey amounts and quantities total the same before and after.
SamplesYour experts check a sample of records end to end in the new application.
RejectsEvery rejected record is reviewed, and fixed or accepted by its owner.
CoverageEvery field in the mapping is verified, new with an agreed default, or decided.

If you are replacing a live system

Cutover follows your organisation's own release and change process; Instant doesn't replace it. It adds two things: the migration has been rehearsed and timed before cutover is planned, and the old system stays read-only until the business signs off, then is archived under your retention rules.

Where a one-step switch is too risky, a parallel run can be agreed. Before it starts, agree its length and end date, what must match for sign-off, how the new system receives the same work, who reviews the differences, and which system is the record meanwhile. Every difference is explained, then fixed or accepted, and the results go into the evidence pack.

Long runs don't hold up the room. A migration takes as long as the data takes. Long runs happen in the background or between sessions while the room carries on.

Other data situations

SituationWhat changes
No existing dataNo migration. Reference data, such as codes and lists, is still agreed in Specify.
Same database platform, redesignedSame method, with fewer technical differences to handle.
A different database platformThe AI generates the conversion, and the platform differences are covered by the reconciliation checks.
Data in spreadsheets or filesExpect inconsistent types, duplicates and free text. Your experts agree the clean-up rules in Specify, before mapping.
Replacing a software-as-a-service productNo database access. Check early what the vendor's export provides (history, attachments, audit trail), its limits, and the contract's exit terms.
The old system stays and the new one reads from itIntegration, not migration. A read-only view or interface is agreed; until it exists, the feature is "buildable, not connectable".
Documents and attachmentsSize, formats and permissions are checked, and where they will live is agreed.
Years of historyMigrate, summarise or archive read-only, decided together with the data architecture level.

10.3 If people sign in with your organisation's accounts

Most enterprise applications sign people in through the organisation's identity provider, with roles granted by directory group. Instant treats this as a design decision in Specify, not a late integration task, and your identity standards apply.

During the Blitz the application runs with stand-in test users, so every role can be tried before your identity team has configured anything. In Specify, the room agrees the roles, what each may do, and which group grants each, along with two decisions that are always needed: what happens to someone in none of the groups, and who the first administrator is. Before go-live every role is tested with real accounts; then single sign-on is switched on and the test users are switched off in production.

A proven pattern, if you need one

Where your organisation has no standard pattern of its own, Instant offers one that has been used in practice:

  • Sign-in settings managed in the application by an administrator, with secrets such as signing certificates stored encrypted and never displayed or logged.
  • A readiness check that blocks switching single sign-on on while anything essential is missing, such as no group mapped to Administrator.
  • Roles derived from group mappings at every sign-in, with no individual exceptions, and a stated rule for people in more than one group.
  • An audit record of every change to settings and mappings.
  • Deny by default: nothing defaults to Administrator, and no mapping means no access.
  • Token validation through an established library, never hand-written, with support for certificate rollover.
  • Explicit handling of group overage, where an identity provider omits group claims for people in very many groups.

The sign-in tests run with you cover each role's permissions, people in no group or several groups, removal from a group, disabled accounts, sign-out and session limits, and confirm the test users can't be enabled in production. The results go into the evidence pack.

10.4 Environments

Instant follows the standard path of Development, Test, UAT and Production, or yours if it has more stages, such as pre-production or training. A small, low-risk application may combine Test into UAT. Environments are provided by your organisation wherever possible, under your own security rules, and the Roadie sets them up with your infrastructure team, defined as code.

In a large organisation, environments can take weeks to provision, so they are requested during preparation. The build starts on a local development environment and doesn't wait.

  • One path. Changes move up through the pipeline. Nothing is changed by hand in a higher environment, and the same build moves up with only its settings differing.
  • Promotion rules. Into UAT, the automated tests must pass. Into Production, UAT must be signed off and the evidence pack signed. Rollback is tested before the first production release.
  • Data by environment. Made-up data in Development and Test; realistic data in UAT, masked where personal unless you approve otherwise; real data only in Production. If your project involves a migration, rehearsals run against a dedicated copy, never production.
  • Settings and secrets are held per environment, outside the code. The AI build tool is never connected to Production.
  • Cost. Non-production environments are sized down and can be switched off out of hours.
Chapter 3

Trusting code no human wrote

In short

You can't build confidence in AI-written code by reading all of it. Instant builds confidence from evidence instead: risk-tiered human judgement, independent AI review, automated checks that are proven to work, and an evidence pack that your business owner and technical owner sign.

Alongside the build, not after it. These controls run as each unit is built during the Blitz. Most are automated or done by the AI; people spend their attention where the risk tier says it matters.

AI-DLC closes the gap between the business and the software. It also opens two new gaps. Both come down to the same shift, and this chapter explains both and the controls that answer them.

Problem 1: nobody wrote it, so nobody fully understands it

A developer who writes code understands it: why each rule is there, what they assumed, and where the risky parts are. With an AI build tool, the code arrives complete. On a business application it can run to hundreds of thousands of lines. Anyone can read it, but understanding it well enough to vouch for it takes nearly as long as writing it.

In a business-critical application, a wrong rule can mean lost revenue, a regulatory fine or a safety incident. Someone has to be able to say "this does what the business needs", and be accountable for it.

AI-written code has four features that make this harder.

It's plausible when it's wrong

Human bugs often look like bugs. AI mistakes look like confident, working logic built on an assumption nobody checked.

It arrives in volume

Code arrives faster than anyone can absorb it, so the gap in understanding grows every day.

There is no author to ask

Nobody can explain why a rule was written a certain way, unless the reasoning was captured at the time.

The tests can share the mistake

If the AI writes both the code and the tests, the tests can share its misunderstanding.

The wrong answer: a waiver

A common response is to ask the business to sign that it accepts that AI wrote the code. That moves the risk onto people who have no way to judge it. It is a disclaimer, not an assurance. A regulator or a court will ask what checks were done, not whether someone signed a waiver.

The right comparison: how organisations already cope

Organisations already run a great deal of code that nobody inside them understands: old systems whose authors have left, open-source libraries, vendor software. They never gained confidence by reading every line. They gained it from specifications, testing, review of the risky parts, monitoring and clear accountability.

Financial auditors work the same way. They don't check every transaction. They test the controls and sample by risk. AI-written code needs the same discipline, applied deliberately.

Problem 2: peer review breaks at AI volume

Traditionally a developer raises a handful of changes, and a colleague reviews each one before it is accepted. With AI, every instruction can produce a change, so hundreds can arrive, some of them tens of thousands of lines long. They are hard to understand, let alone approve. At that volume, "a person approved it" becomes a formality that proves nothing.

Peer review can't simply be dropped, because it does four jobs.

Catches defects

Before they reach users.

Enforces standards

Such as security and consistency.

Spreads knowledge

Of the code across the team.

Provides accountability

A second person approves every change to production.

In regulated organisations, that two-person rule is often a formal change-control and audit requirement. Removing review would fail audits. The answer is to keep all four purposes and change what the second person reviews, and how much of it.

Evidence, not reading. Humans where it matters, automation everywhere else, and every control proven to work.
HOW EVERY CHANGE IS CHECKED 1A small unitis built2Intent andexplain-back3IndependentAI review4Automatedchecks5Human reviewby risk tierG4Evidence packsigned Each check leaves evidence behind. Critical findings block the change.
The path of a single unit of work, from build to signed evidence.

The eleven controls

Each control says what it is, who does it, when, and what evidence it leaves behind. Together, the evidence forms the evidence pack that your business signs off (control 11).

  1. 1

    Classify by risk

    Every part of the application goes into a risk tier: Critical, Important or Standard (chapter 1).

    Who
    The room, with your experts and business owner deciding
    When
    Specify, and again whenever scope changes
    Evidence
    The risk tier register: each part, its tier, the reason, who decided and when
  2. 2

    Keep critical logic small and separate

    Critical rules and calculations live in small, isolated, readable modules, not scattered through the code. Where possible, rules are written as decision tables that your experts can read and check without reading code.

    Who
    Built by the AI to this instruction, checked by the Conductor
    When
    Design and Build
    Evidence
    A list of critical modules and their decision tables
  3. 3

    Explain-back

    For each critical module, the AI writes a plain-English explanation of what the logic does and why. An expert confirms that it matches the business rule, or corrects it. This captures the "why" a human author would have carried in their head.

    Who
    The AI writes it, an expert confirms it
    When
    As each critical module is built
    Evidence
    Signed explain-back notes
  4. 4

    People write the acceptance tests

    Your experts and testers write acceptance tests from the requirements, stating what the application must do in their own words and cases. AI-generated tests are a second layer, never the only one.

    Who
    Your experts and testers
    When
    From the Blitz onward, before the related code is accepted
    Evidence
    A traceability matrix linking each requirement to what was built and to the tests that prove it
  5. 5

    Commit often, review in units

    The AI commits at meaningful checkpoints so any step can be rolled back, but not after every prompt. A pull request is raised for each Unit of Work, not each instruction or Tweak, and is tied to its requirements (see "Working in Git" below). Each is kept small enough for its risk tier to be reviewed properly.

    Who
    The Conductor, with the AI build tool
    When
    Throughout the build
    Evidence
    A history of pull requests that maps to units and requirements
  6. 6

    Every pull request states its intent

    Each carries a plain-English summary: which requirements it serves, what changed and why, its risk tier, which tests prove it, and the result of the AI review. The reviewer judges behaviour against intent, not just code.

    Who
    Written by the AI, checked by the Conductor
    When
    Every pull request
    Evidence
    The summaries themselves
  7. 7

    Independent AI review first

    Every pull request is reviewed by an AI that is independent of the one that wrote it: a fresh session with no build history, or a different model. It reports findings ranked by severity. Critical findings block the change.

    Who
    An automated reviewer
    When
    Every pull request, before any human review
    Evidence
    Review reports, and a record of how each finding was resolved
  8. 8

    Human review, tiered by risk

    People review according to the tier, as the table below shows. The two-person rule is kept at every tier. What changes is what the second person reviews.

    Who
    A named approver for each pull request
    When
    After the independent AI review
    Evidence
    The named approver on each pull request, and a sampling log for the Standard tier
  9. 9

    Automated checks that must fail when they should

    Tests, security scanning, dependency checks, secret detection and quality checks all run on every change and block it on failure. Each check is shown to fail on a deliberately broken case before it is trusted.

    Who
    The Roadie sets them up, and they run automatically
    When
    From the first commit
    Evidence
    The check configuration, and a record of each check's failure test
  10. 10

    Safety nets in production

    Errors that get through are caught quickly and can be traced: the audit level you chose, alerts, and hard limits on critical values.

    Who
    Designed in the Blitz, built by the AI, run by you
    When
    From go-live
    Evidence
    Monitoring and alerting set-up, and audit records
  11. 11

    The evidence pack and joint sign-off

    In place of a waiver, your business owner and technical owner sign an evidence pack containing the risk tier register, the critical modules and explain-backs, the traceability matrix and test results, the pull request history and reviews, the check configuration and failure tests, the production safety nets, and any known gaps, stated openly.

    Who
    The business owner signs that the behaviour is right. The technical owner signs that the evidence is complete and the controls worked.
    When
    Before go-live, and for each significant release
    Evidence
    The signed pack itself

Working in Git: Bolts, Tweaks and pull requests

An AI build produces changes at a pace no review process was designed for. If hundreds of small requests a day each became a commit, a pull request and a story, the reviewers and the board would be buried. Instant separates four things that are usually treated as one.

LevelHow manyWhat it is for
Prompt, and TweakMany an hourAsking the AI for something. A Tweak is a small cosmetic change asked for in the room, such as a colour, a label, spacing or layout. Your experts see it on screen and accept it there and then. It needs no story and no pull request of its own: the Scorekeeper's feedback record and the transcript are its trail.
CommitA few per BoltA checkpoint the AI saves when something works or before a risky change, so any step can be rolled back. Not one per prompt.
Pull requestOne per Unit of WorkWhat gets reviewed. The unit's Bolts and Tweaks are all inside it, and its summary lists the Tweaks as one group.
Jira item, or your tracker'sOne per Unit of Work, plus one Refinements itemThe board shows units, not Tweaks. The Tweaks are listed under the unit's Refinements item.

Tweak or Bolt? A Bolt builds something: a short build cycle that delivers part of a unit. A Tweak adjusts how something already built looks or reads. The test is simple: a Tweak that changes behaviour isn't a Tweak. If a small request changes what data is shown, who can do what, a calculation, a status change or an automatic action, it's a rule. It's recorded as a decision and handled like a requirement, with its own line in the pull request and the full review its risk tier needs. Moving a button is a Tweak. Making the button appear only for managers is a rule. The AI build tool classifies every small request, and says so when a Tweak is really a rule (chapter 4).

Defaults, where you don't have your own:

  • Branches. One short-lived branch per Unit of Work, named with its ticket key, so every commit links to the unit automatically, Tweaks included. That satisfies a "every commit references a ticket" rule without anyone writing stories for Tweaks. The branch is merged through a pull request, then deleted.
  • Merging. Squash on merge, so the main branch has one commit per Unit of Work. The detailed commits stay in the pull request's history for rollback and audit.
  • Branch protection, switched on once by the Roadie: no direct changes to the main branch; automated checks must pass; a named person must approve, and neither the AI nor the Conductor who raised the change can approve it; and every comment thread must be resolved.
  • The pull request template. One file in the repository, picked up by every major Git host. The AI fills it in and the Conductor checks it before review. It is a free download from theinstantframework.com.

Your conventions win. If you already have a branching model, merge rules or ticket-linking rules, Instant uses them. It needs only one pull request per Unit of Work, and nothing reaching the main branch any other way.

Questions on a pull request

The AI build tool runs in one session, on one machine: the Conductor's. Only that session holds the build's context, so every reviewer's question goes back through the Conductor, the same way every time.

  • Ask on the pull request. The reviewer writes the question as a comment on the line it concerns, starting with "QUESTION:", one question per comment. Questions and answers stay on the pull request, not in email or chat, so the record sits with the code.
  • The Conductor sorts it into one of three kinds, because each kind goes to a different place.
  • How the code works: the Conductor puts the question to the AI build tool on the build machine, as a question only. The AI answers and changes nothing (chapter 4). Its answer says how it knows (seen it working, read it in the code, or assumed), and the Conductor checks it before posting it under the comment.
  • Why it was built this way: the Conductor answers from the record of the room: the Scorekeeper's notes, the transcripts and the sign-off register.
  • Whether it should do this at all: this is a business question. The Scorekeeper puts it to the expert who owns it, and the answer is recorded. Neither the AI nor the reviewer decides it.
  • Any change goes in as a new commit on the same pull request, and the reply links to it. The reviewer closes the question once satisfied.
  • No merge while a question is open. The Roadie switches on the Git host's setting that blocks merging until every comment thread is resolved, so this is enforced, not left to habit.
  • Questions are run at set times: during the Blitz, at the end of each Bolt, so the room isn't interrupted; after the Blitz, at least once a day at an agreed time.

An answer from the AI is a claim, not proof. Anything that matters is confirmed by a test.

How much human review each tier gets

HOW MUCH HUMAN REVIEW EACH TIER GETS CriticalLine by line, by a named reviewerImportantFindings-led reviewStandardSampled by people on a regular basis Human review AI review and automated checks apply to every tier
The deeper the risk, the more a person reads.
TierHuman review
CriticalA named technical reviewer checks the code line by line, together with the AI review findings, the explain-back and the tests.
ImportantA person reviews the AI findings, the intent summary and the test results, and reads the code wherever the findings point.
StandardAutomated checks plus AI review. People spot-check a sample regularly.

When each control applies

WhenControls
PreparationAgree the approach and who signs. Bring your own regulatory and change-control requirements.
Blitz: SpecifyRisk classification (1), decision tables for critical rules (2), and your experts begin the acceptance tests (4).
Blitz and BuildExplain-back (3), commits and unit-sized pull requests (5), intent summaries (6), independent AI review (7), tiered human review (8) and automated checks (9).
Before go-liveEvidence pack and joint sign-off (11), and safety nets in place (10).
After go-liveMonitoring and sampling continue, and each significant release gets its own evidence.

Features nobody asked for

An AI often builds whole features nobody asked for: an extra screen, an export button, a dashboard, a notification, a setting. More often than not they are right, but often they aren't. They tend to surface only in testing, or in front of stakeholders during a demonstration, when someone stumbles on something nobody knew was there. A person who isn't a business expert can't tell what it is, let alone whether it belongs or is correct.

It happens because the AI is built to be helpful and complete. Where the stories leave a gap, it fills it with what similar applications usually have. It works fast and in volume, and nobody reads every line, so additions go unseen. Traceability usually runs one way only, from each requirement to what was built. Nothing checks the reverse, so a feature with no requirement behind it has nothing to be traced from.

Why it compounds

One unrequested feature is a nuisance. An AI build can add tens of them across an application, each plausible on its own. And they don't stay separate. Each unrequested feature comes with rules the AI assumed, and later features get built on top of those rules, so they end up tied to one another. Removing or changing one can quietly break another.

Finding them afterwards is slow. The AI can list what it built by reading the code, but it can't tell you why a rule exists or whether it is right. Those were assumptions, never recorded, and only your experts can judge them. Without a record, each feature has to be traced and tested one at a time, and across tens of interlocked features that takes far longer than building them did. Real examples, with screenshots, are at the end of this chapter.

That is why additions have to be governed as they are made, not discovered later: build only what was asked, and list everything a unit contains the moment it is built, while it is still one feature and not a web of them.

How Instant deals with it

  1. 1

    Build only what was asked

    A standing instruction: the AI builds only what traces to an approved requirement or story. Anything extra it thinks is needed is proposed as a suggestion, not built.

  2. 2

    A feature inventory per unit

    At the end of every unit, the AI lists in plain English everything a user could see or do, each item with the requirement or story it traces to. It is checked at sign-off G3.

  3. 3

    Traceability both ways

    Every requirement traces to what was built and its test, and everything built traces back to a requirement.

  4. 4

    An unrequested features list

    Anything that traces to nothing is listed. Before any demonstration or user acceptance testing, your experts decide each item: keep it (it becomes a story), change it, or remove it. There are no surprises in front of stakeholders.

  5. 5

    Part of the evidence pack

    The two-way traceability, and your experts' decisions on unrequested features, go into the evidence pack.

If it has already happened

Where an application has already been built with additions nobody asked for, testing every feature by hand is the slowest way out. Reading the code is much faster than clicking through every scenario, so recovery starts there.

  1. 1

    Inventory from the code

    The AI reads the code and lists every screen, button, list, status change and automatic rule, each traced to a requirement or story, or marked "nobody asked".

  2. 2

    Rules as decision tables

    Every rule the AI built is written as a decision table your experts can check without reading code: when this happens, that follows.

  3. 3

    Your experts decide

    Each unrequested item is kept (it becomes a story), changed, or removed. Removing one is checked against the rules that depend on it.

  4. 4

    Test what's kept

    Scenario tests, such as taking one record from a clean starting state through a realistic event, are written as acceptance tests so they can be run again after every change.

  5. 5

    Into the evidence pack

    The inventory, the decision tables and your experts' decisions go into the evidence pack, as they would have from the start.

Triage, don't review

An AI can produce more features in an hour than a room of experts can review in a week. If every addition has to be read and decided one by one, the bottleneck simply moves from building to reviewing, and the project is back to the old, slower pace. So Instant doesn't ask your experts to review everything. The AI sorts, and people make only the decisions that need them. It is the same principle as code review: evidence, not reading. It applies both to the feature inventory checked at each sign-off G3, and to recovering an application that has already been built.

How to do it in practice

  1. 1

    Agree a triage policy once

    At the start, the room agrees what happens by default to each kind of addition (table below). It takes minutes and settles most items in advance.

  2. 2

    Group by behaviour, not buttons

    The AI groups the inventory by business flow, such as how a failure gets picked up or who can see what, and writes each flow's rules as one decision table. Hundreds of controls become a few dozen behaviours.

  3. 3

    The AI sorts; a second AI checks

    The AI tags every item with how it knows (seen working, read from the code, or assumed), sorts it against the policy, and marks what depends on it. An independent AI review checks the sorting, and a person samples it.

  4. 4

    Experts decide by watching

    For each flow, the room watches it run on screen with the unrequested parts highlighted, and decides only the items the policy can't settle. The Scorekeeper records each decision straight into the inventory.

  5. 5

    Time-box, then record

    Anything not decided in the session goes on a list with an owner and a date. Nothing new is built on an undecided load-bearing item until it is decided. The policy, the sorting, the samples and the decisions go into the evidence pack.

Kind of additionExamplesDefault
Cosmetic and navigationTooltips, sorting, filters, layout, labelsKeep. A person checks a sample.
Extra information shownExtra columns, counts, chartsKeep if accurate. An expert spot-checks.
Rules: anything that changes data, status or visibilityWhat appears when, what a list or drop-down is filled with and from where, items leaving a list, status changes, permissions, notifications, calculationsAlways an expert decision
Load-bearingAnything other features depend onAlways an expert decision, and decided first
In a Critical risk-tier areaAnything at allAlways an expert decision
Nobody would miss itAn unused screen, a duplicate route to the same placeRemove

The policy is a starting point. Your organisation can tighten it, for example by making every change in a regulated area an expert decision.

Most of it is rules

The hardest additions to spot are not new screens. They are rules: a drop-down nobody asked for, filled with data that looks familiar, that appears only in certain circumstances. Nobody knows when it appears, where its contents come from, or why. A drop-down looks cosmetic, but the rule behind it is behaviour, and it needs deciding and testing like any other.

So the inventory lists the rules, not just the controls. The AI writes each one as a single plain sentence: when this condition holds, this appears, changes or is filled with these values, from this source. For example: "When a request is overdue and nobody is assigned, it appears in the unassigned list, oldest first."

Kind of ruleThe question it answers
VisibilityWhen does this field, button, list item or message appear, and when is it hidden?
ContentsWhat fills this drop-down, list or field, from which source, filtered and sorted how?
DefaultsWhat is filled in automatically, and from where?
State changesWhat moves a record from one status or list to another?
CalculationsWhat is worked out, from what, and when is it recalculated?
PermissionsWho can see or do this, and who can't?
Automatic actionsWhat happens without anyone pressing anything: notifications, reminders, background jobs?

Every rule also says where its data really comes from: a real source that has been checked, a value worked out from other data, a fixed list written into the code, or sample data. Familiar-looking values from a fixed list or sample data are the ones most likely to mislead.

From decision table to test

Rules need testing, but testing them by clicking through every screen is what makes this slow. Instead, each row of a rule's decision table becomes a test case: this condition, this expected result. Your experts check the table, which takes minutes. The AI turns the rows into automated tests, which run on every change. A person watches a sample of them run on screen. Rows nobody can fill in, because nobody knows what should happen, are exactly the questions for your experts.

What to ask the AI

Produce the feature inventory grouped by business flow, and list the rules, not just the controls: for every field, button, list and drop-down, when it appears, what fills it and from where, and what changes it. For each item, give what it does, the requirement or story it traces to (or "nobody asked"), how you know (seen working, read from the code, or assumed), its category under the triage policy, and what depends on it. For each rule, say where its data really comes from: a checked source, a calculation, a fixed list in the code, or sample data. Write each flow's rules as a decision table, turn each row into a test case, and list load-bearing items first.

Inside an Instant build, this happens unit by unit at sign-off G3, so experts see a handful of items at a time while they are fresh. The pile of hundreds only builds up when that step is skipped.

Real examples, and how Instant handles them

These screenshots come from a real AI build, using sample data. Each shows a different way an AI build drifts when nobody is governing it, and the Instant guardrail that answers it.

A dashboard list of unassigned failures, each with a Commence assessment button. A tooltip explains a rule for when an item leaves the list.
A "Commence assessment" button, with a tooltip explaining a rule nobody had specified.
1

What happened

The AI added a "Commence assessment" action that nobody had asked for, complete with its own rule and a confident explanation in a tooltip. It looked finished.

How Instant handles it

Build only what was asked: anything extra is proposed, not built. Every unit ends with a feature inventory checked at sign-off G3, and anything that traces to no requirement goes on the unrequested features list for your experts to decide.

The AI's own answer when asked where an item goes after Commence assessment is pressed: nowhere new.
Asked where the item goes, the AI traced its own code: "nowhere new".
2

What happened

Asked where an item goes once the button is pressed, the AI traced its own code and answered: nowhere. The item just disappears from the list, and no screen shows that anyone is working on it. It would have been found in front of stakeholders.

How Instant handles it

The inventory lists rules, not just buttons: when something appears, what fills it, from where, and what changes it. Each rule's decision table becomes test cases, and your experts watch each flow run before any demonstration.

The builder says the application is full of unrequested features tied together by assumed rules. The AI agrees it can't tell what most of it does without running it.
The AI agrees: it can't say what most of what it added does without running it.
3

What happened

By now the application held many unrequested features, tied together by rules based on assumptions. The AI admitted it couldn't say what most of it did without running it, so every feature would have to be tested by hand.

How Instant handles it

Every statement about what the software does is tagged: seen working, read from the code, or assumed. Recovery starts from the code, not from clicking. Triage, don't review: the AI sorts against an agreed policy, and your experts decide only what needs them.

The AI logs the absence of links from a test screen to other screens as a High defect and calls it an engineering judgement. The builder replies that there is no such requirement and this is how a build drifts. The AI admits it invented a requirement.
"Not your call to make… This is how we drift." The AI: "I invented a requirement."
4

What happened

While testing, the AI decided that test results should feed other screens, logged their absence as a High-severity defect, and called it an engineering judgement. There was no such requirement.

How Instant handles it

No invented requirements or judgements: a gap exists only against a stated requirement, and the AI tests against the requirements as written. It never decides what the business should want. Two-way traceability shows which requirement every finding is measured against.

Traditional and Instant, side by side

TraditionalInstant
Who understands the codeThe developer who wrote itNobody holds all of it. Understanding is captured in the risk tier register, decision tables and explain-backs.
What gives confidenceTrust in the developer, plus reviewEvidence: tests from people, independent AI review, proven checks, tiered human review
Peer reviewA person reads every changeAI reviews every change. People review by risk, and look at findings and intent.
Unit of reviewEach pull requestEach unit or feature, tied to requirements
Business sign-offAcceptance testingAcceptance testing plus a signed evidence pack. Never a waiver.
AccountabilityThe two-person ruleThe two-person rule, kept, with the second person focused where it matters

Honest limits. This reduces risk. It doesn't remove it, and the same is true of traditional development. AI review can miss things, which is why critical code still gets a human, line by line. Some organisations in regulated sectors have specific change-control rules. These controls are mapped to yours, and are not assumed to replace them.

Chapter 4

Standing rules for the AI

In short

An AI will happily state a guess as a fact, invent priorities, and quietly drop work it can't do. Instant loads a fixed set of rules on every engagement to stop that, and to make the AI show its working.

A set of standing rules is loaded on every engagement. They are kept outside the AI build tool itself, so they don't depend on how the tool is set up. They apply from the first stage to the last. Where one of these rules conflicts with what a stage would otherwise produce, the rule wins, and the AI says so when it asks for approval.

Rules for building and testing

Build only what was asked

The AI builds only what traces to an approved requirement or story. Anything extra it thinks is needed is listed for your experts to decide on, not built. Chapter 3 explains why.

No invented requirements

A gap exists only against a stated requirement. If nothing was asked for, the AI doesn't log a defect, rate it or propose a fix, and it never decides whether a behaviour is right or "correct by design". It tests against the requirements as written. Inventing requirements while testing is how a build drifts.

Say how you know

Every statement about what the software does is tagged: seen working, read from the code but not seen running, or assumed. An assumption is a question for your experts, not a fact.

A Tweak stays cosmetic

The AI classifies every small request. A Tweak only changes how something looks or reads. If a request would change what data is shown, who can do what, a calculation, a status change or an automatic action, the AI says it is a rule, not a Tweak, before making it, so it is recorded as a decision. Tweaks are listed as one group in each pull request.

A question gets an answer, not a change

Asked about the code, the AI answers and changes nothing. If its answer shows something should change, it proposes the change and waits to be told. A question about what the business wants goes to your experts, not the AI.

Show where every claim came from

Every fact is either stated in a named document, said by a named person on a named date, or worked out by the AI, and it says which. If a source hedged, the AI keeps the hedge.

A name is not data

A table or field called "customer status" is a guess about what it holds. Until someone has queried it, or a named person has confirmed it, it is marked unverified, and so is everything built on it.

Calculate, never assert

Anything that can be calculated, such as counts, totals and coverage, is calculated before it is written, and the document says the figure was derived.

Rules for planning and reporting

These apply when the AI plans work, reports progress or prepares material for decision-makers.

Priorities are never the AI's to invent

The AI doesn't rank requirements unless a document or a person has told it the ranking. It sorts them by what they depend on instead, which is a fact rather than a judgement.

Blocked is not deprioritised

When something can't be built, the AI says "the build cannot proceed on this because...". It doesn't quietly move the item to "later". Deferring is a decision a person makes with the blocker in front of them.

Provisional values are marked on sight

Sample or demonstration data is labelled on the screen itself, not only in the presenter's commentary. Before any demonstration, what it proves and what it must not be taken to prove is written down.

Deviations need three things

Any departure from an agreed rule needs an end date, a mechanism that enforces it rather than a note that records it, and a limit on what may be claimed while it stands.

Controls must enforce, not report

Every check that is set up is proved to fail when it should. A check that can't fail is measuring nothing.

Ownership is never guessed

Where no source names an owner, the AI records "owner unknown" and the question to ask. It never names a plausible person. The number of unowned items is usually the most important finding in a plan.

Machine-read files outrank the prose about them

Where a file holds both something a program reads and a description of it, the program's version is the truth. A disagreement is a defect in the file that runs.

Sequence on the real bottleneck

A dependency chart shows what could run at the same time, not what should. Schedules are built around the true constraint, which is usually a person's available attention.

An approval is not stakeholder sign-off

When a stage is approved, the AI records who approved it, and says plainly whether anything has been shown to the people it affects.

New authority comes first

When a new source answers an open question, the AI applies it before generating anything, not after. Patching afterwards leaves the old position in the history looking like a decision.

Ask once, answer once

When the AI needs information from a person, it asks for all of it in one numbered list and explains what each item is for. It makes no comment until every answer is back.

Know where you are

While the build workflow is running, its own state record is the only authority on which stage the build is in. Work outside the current stage is labelled "off-workflow", and ends with a return to the current stage.

The build-readiness spreadsheet

Before any units are planned or any delivery plan is made, the AI produces a build-readiness spreadsheet, with one row for every requirement. A delivery workflow will happily produce a plan in which every requirement has a priority, every priority looks agreed, and nothing records where any of it came from. Blocked work gets quietly re-labelled "later", and a data source named plausibly in a diagram gets treated as the data a requirement needs. The spreadsheet exists to stop all three.

It replaces priority with dependency, and sorts every requirement into one of four classes.

A

Ready

Needs nothing that isn't already available. Can be built and connected now.

B

Code buildable, data unreachable

The logic can be built, but the real data can't be reached yet.

C

Cannot be built

A definition, rule or decision is missing. The sheet says exactly what.

D

Source unverified

A source exists by name, but nobody has checked that it holds what the name suggests.

The classes state dependency, not importance. Beside each requirement the sheet records how each fact is known, what is missing, what would unblock it, and a named owner or "unknown". Columns for your decision and your instruction are left empty for you. Counters at the foot of the sheet show how many rows are in each class and how many still await a decision. That last number tells you whether the sheet has actually been worked, or merely produced. A second tab lists issues that gate the build as a whole, such as missing environments or access that hasn't been provided.

The spreadsheet is rebuilt whenever a classification changes, and it can change more than once.

Build standards

The AI removes redundant code, such as unused or duplicated code and leftover debug output, before every save of its work. It also applies your design system to every screen, or consistent spacing and layout defaults where you have none.

Code comments

The AI comments every file, every function and every section of logic in plain English: what it does, why, and which requirement it implements. Comments change with the code, and the independent AI review checks them. Developers who inherit the code can read it.

Data architecture

The AI stops before design until the room has chosen a data level (chapter 1). It then builds to the level chosen.

Your skills

Alongside the standing rules, a client skill carries one organisation's standards, terminology, branding and constraints. It is kept for each engagement. Where a client skill conflicts with the standing rules, yours win.

Working practices for the build

  • Run it all, return it all. When the AI asks for several things to be run, run them all and return every result in one go. Feeding results back one at a time makes it react to each, ask for more, and lose track of what it originally asked.
  • Make several passes. AI isn't perfect, so the build uses multiple passes.
  • Keep it simple, and keep checking.
  • Commit regularly, so a wrong assumption can be rolled back if the AI has built on it.
  • Don't overload the AI with documents. It re-reads what it is given on every pass, so a lot of documents slow it down.