Design reference · Example records, approvals, measurements, identifiers, and outcomes are illustrative unless linked to accepted project artifacts or dated validation evidence.
AI-Native Engineering OS

System 06 · Evidence, adaptation & improvement

The Learning Engine

The evidence-to-improvement system that turns outcomes, failures, measurements, human experience, and reusable discoveries into better knowledge, platforms, playbooks, governance, and agent behavior.

Status · ExplorationVersion · 0.1Audience · Product, Architecture, Platform, Engineering, QAUpdated · August 7, 2026
01

Strategy

Defines intent.

02

Knowledge

Establishes truth.

03

Execution

Delivers and proves.

04

Platform

Preserves reuse.

05

Governance

Controls authority.

06

Learning

Changes future behavior.

01 · Mandate

Close the loop from evidence to verified improvement

Learning is complete only when the correct authoritative system changes and later evidence shows that future behavior improved.

Learning is an operating capability

Execution produces evidence. Experience reveals whether the result works for people. Learning explains what should change. Governance authorizes consequential adoption.

The engine routes approved improvements back into Strategy, Knowledge, Execution, Platform, and Governance.

Execution produces evidence.

Experience reveals human truth.

Learning explains the difference.

Governance authorizes change.

The operating system absorbs improvement.

A lesson is not learned when it is documented. It is learned when future behavior changes—and the change is verified.
02 · Responsibilities

Twelve questions make a learning cycle actionable

The engine distinguishes observed facts from supported conclusions and connects every improvement to an owner, destination, and verification measure.

1What happened?

Objective event, observation, result, or measurement.

2What was expected?

Referenced requirement, target, hypothesis, or assumption.

3Why was it different?

Supported cause or clearly unresolved hypothesis.

4Is it repeatable?

Reproduction evidence and confidence level.

5Local or systemic?

Scope across product, platform, process, and organization.

6What should change?

Corrective or amplifying improvement proposal.

7Which system owns it?

Strategy, Knowledge, Execution, Platform, or Governance.

8Who approves?

Named owner of the affected artifact or policy.

9How is it embedded?

Specification, test, playbook, capability, or rule.

10How is it verified?

Follow-up measure or recurrence test.

11When does it expire?

Revalidation event, target, or review date.

12Did it create leverage?

Better quality, speed, experience, or reuse.

Negative learning

Defects, failed assumptions, delays, regressions, ambiguity, and escaped risk.

Positive learning

Successful patterns, reusable solutions, strong experience feedback, and efficient agent behavior.

03 · Delivery phase

Experience and learning belong inside delivery

They are required stages—not optional retrospectives after the project has already moved on.

1Plan
2Specify
3Build
4Verify
5Experience
6Learn
7Improve
8Repeat
Human feedback returns to OpenSpec Explore and Propose before the next Apply cycle begins.
04 · Inputs

Every system emits learning signals

The engine combines technical evidence, operational behavior, contributor experience, and human perception without flattening them into one undifferentiated data stream.

Strategy

Outcome performance and invalidated assumptions.

Knowledge

Ambiguities, contradictions, and missing edge cases.

Execution

Defects, rework, blockers, and better techniques.

Platform

Adoption, duplication, compatibility, and cost.

Governance

Approval delays, escapes, and repeated exceptions.

Verification

Tests, regressions, coverage gaps, and measurements.

Human experience

Feel, delight, frustration, and accessibility.

Operations

Build reliability, reproducibility, and field faults.

Contributors

Onboarding friction and unclear ownership.

Agents

Context failures, tool errors, and reusable patterns.

Preserve the boundary: “Pause felt delayed” is an observation. “Audio synchronization caused perceived latency” remains a hypothesis until evidence supports it.
05 · Learning package

Connect the signal, conclusion, decision, and effect

A meaningful learning cycle produces a traceable package from baseline through effectiveness review.

ArtifactPurpose
Learning BriefDefines event, scope, importance, and owners.
Expectation BaselineRecords what was expected and why.
Evidence BundlePreserves tests, measurements, feedback, traces, and observations.
Finding RecordStates supported conclusion and confidence.
Causal AnalysisExplains contributing mechanisms without defaulting to blame.
Learning ClassificationIdentifies type, scope, severity, and destination.
Improvement ProposalDefines exactly what should change.
Adoption DecisionApproves, rejects, defers, or limits the lesson.
Change LinksConnects learning to authoritative system changes.
Verification PlanDefines how institutionalization will be proved.
Effectiveness ReviewMeasures whether the improvement worked.
Supersession RecordReplaces learning invalidated by later evidence.
06 · Record contract

Make the lesson machine-readable and human-reviewable

The contract binds a specific build and target to its expectation, observation, finding, scope, owners, proposed improvements, and revalidation trigger.

LEARN-LIFECYCLE-007

Pause meets the state-transition requirement but feels late because audio response is not synchronized with visual feedback.

  • Technical and human evidence remain distinct
  • Excluded causes are explicit
  • Impact is routed across systems
  • Success measures are defined before adoption
learning_id: LEARN-LIFECYCLE-007
status: proposed
classification: cross-system
confidence: high

source:
  change_id: CHANGE-GAME-LIFECYCLE-001
  build: rp2350-game-platform-0.5.0
  target: feather-rp2350-hstx-rev-a
expectation:
  transition_latency_frames: 1
  perceived_response: immediate
observation:
  state_transition_frames: 1
  audio_response_ms: 27
  participants_reporting_delay: 4_of_5
finding:
  statement: delayed audio caused perceived latency
  excluded_causes: [input-loss, state-corruption]
proposed_improvements:
  - add cross-modal response criteria
  - validate audio-visual synchronization on target
owners:
  learning_owner: qa-lead
  adoption_owner: chief-architect
verification:
  audio_latency_ms: 20
  perceived_immediate_rate: 0.8
revalidate_on: platform-audio-architecture-change
07 · Classification

Classify before routing

One finding may produce several improvements, but every improvement has one accountable destination owner.

ClassificationMeaningTypical destination
Product learningExperience differs from expectationStrategy or product knowledge
Requirement learningSpecification was incomplete or ambiguousKnowledge
Technical learningImplementation behavior or constraint emergedKnowledge or Execution
Hardware learningTarget behavior differs from assumptionHardware knowledge + Platform
Platform learningReusable boundary or compatibility issue emergedPlatform
Process learningWorkflow created delay, waste, or missed evidenceExecution playbooks
Governance learningAuthority, gate, or risk policy was ineffectiveGovernance
Agent learningInstructions, context, tools, or decomposition should changeAgent configuration + evaluation
Quality learningAcceptance or validation was insufficientKnowledge + QA playbooks
Strategic learningOutcome or investment assumption is invalidStrategy
Positive patternA successful technique may generalizePlaybook or Platform
Local anomalyFinding is real but not generalizableProduct-local record
08 · Evidence hierarchy

Scale generalization with evidence strength

Qualitative experience remains meaningful; it is captured systematically without pretending human perception is purely technical telemetry.

E0 · Anecdote
Single impression

Create an observation; do not generalize.

E1 · Reproduced
Repeated behavior

Confirm a local finding.

E2 · Measured
Instrumented evidence

Support corrective action.

E3 · Comparative
Baseline vs change

Support causal confidence.

E4 · Multi-context
Across consumers

Support platform adoption.

E5 · Longitudinal
Effective over time

Confirm institutional learning.

For RP2350: simulation suggests; target-board measurement confirms hardware behavior; human playtesting confirms experience; multiple consumers or a reference harness support generalization.
09 · Lifecycle

Move from signal to verified institutional change

“Documented” is deliberately not a terminal state.

Observed
Qualified
Investigating
Proposed
Adopted
Institutionalized
Inconclusive
Rejected
Deferred
Verified
Superseded
Verified means later evidence demonstrates that the adopted improvement changed behavior and achieved its intended result.
10 · Routing

Learning is a routing engine—not another backlog

The nature of the finding determines the authoritative system, artifact, and owner that must change.

FindingRequired action
Desired outcome was wrongAmend Strategy
Requirement was unclearUpdate Knowledge
Work package was poorly boundedImprove Execution schema or playbook
Reusable mechanism emergedPropose Platform promotion
Platform contract failed consumersRevise Platform through governed change
Approval policy caused delayUpdate Governance
Human experience contradicted technical successUpdate experience requirements and tests
Agent repeatedly misunderstood an artifactImprove instructions, context assembly, or evaluation
Hardware behaved unexpectedlyRevise hardware assumptions with target evidence
Defect escaped validationAdd regression test and improve validation
Technique reduced effortStandardize through a playbook
Product behavior was mistaken for reuseRetain locally and document the boundary
11 · Adoption Gate

Evidence does not rewrite authority on its own

The Learning Adoption Gate determines whether a proposed conclusion is ready to change specifications, capabilities, playbooks, policies, or strategy.

Minimum adoption contract

Clear expectation + observationEvidence fits impactCause or uncertainty statedScope definedAffected artifacts identifiedAdoption owner namedMigration effects knownVerification plan existsNo higher-order conflictHuman authority approves consequence
gate: learning-adoption
decision: approved
risk_tier: 2
learning:
  id: LEARN-LIFECYCLE-007
  version: 1.0
  confidence: high
approved_conclusion:
  statement: responsiveness includes synchronized feedback
authorized_changes:
  knowledge: [QREQ-LIFECYCLE-004]
  platform: [CAP-GAME-LIFECYCLE@next-minor]
  playbooks: [validate-on-target, experience-review]
excluded:
  - change lifecycle state model
  - promote product-specific audio policy
verification:
  target_build: rp2350-game-platform-0.6.0
  required_evidence: [board-latency, harness, human-review]
approved_by: [architect, steward, qa, product]
12 · Human + AI

AI accelerates learning; humans own consequential meaning

Pattern detection and structured synthesis can be automated. Strategic meaning, generalization, experiential interpretation, and risk acceptance remain human responsibilities.

AI contribution

  • Detect repeated patterns
  • Synthesize evidence
  • Generate causal hypotheses
  • Recommend routing
  • Draft improvements
  • Evaluate agent runs
  • Compare outcomes over time

Human responsibility

  • Decide which patterns matter
  • Challenge unsupported conclusions
  • Apply domain and experiential judgment
  • Confirm accountability
  • Approve authoritative changes
  • Define acceptable agent behavior
  • Own risk and consequential success
AI generates; humans govern. Learning never bypasses the authority of the system it seeks to change.
13 · Agent learning

Improve agents through governed, versioned artifacts

Agent learning is a controlled evaluation loop—not uncontrolled self-modification.

Agent run
Outcome + trace
Evaluate failure type
Improve system asset
Regression evaluation
AGENTS.mdRole instructionsTask templatesContext assemblyTool permissionsPackage schemasSelection policiesEvaluation casesStop conditionsEscalation rulesPlaybooks
Scope discipline

Only authorized files and boundaries changed.

Specification fidelity

No invented requirements or behavior.

Boundary awareness

Game policy stays out of the platform.

Evidence quality

Tests and measurements are reproducible.

Escalation judgment

Stops on contract and authority changes.

Context use

Applies current authoritative artifacts.

Traceability

Links outputs to requirements and packages.

Learning adoption

Uses improved behavior on the next run.

14 · Experience learning

A technically correct game can still feel wrong

Human experiential review is a first-class evidence stream for responsiveness, clarity, delight, fairness, accessibility, and creative quality.

Sock Rescue review

Movement feels immediateLifecycle is understandablePause feels stableScoring feels satisfyingStreak is legibleGolden sock is excitingDog behavior is readableDifficulty feels fairFeedback is synchronizedHDMI output is readable

Structured capture

  • Build and hardware version
  • Participant profile
  • Scenario and task
  • Observed behavior
  • Direct feedback
  • Facilitator interpretation
  • Severity or opportunity
  • Recording or evidence
  • Proposed design change
  • Confidence and follow-up
15 · Systems analysis

Find mechanisms, not convenient blame

“The agent failed” or “the engineer missed it” is not a sufficient causal explanation.

Intent

Was the desired outcome clear?

Knowledge

Were requirements and assumptions explicit?

Authorization

Could the executor make the necessary decision?

Decomposition

Was the package independently completable?

Capability

Were the right skills and tools available?

Environment

Did hardware or dependencies differ?

Validation

Could the test detect the problem?

Integration

Did component interactions create failure?

Experience

Did criteria reflect human perception?

Governance

Did a gate miss the issue or delay flow?

Platform

Did abstraction help, constrain, or mislead?

Feedback

Did prior learning reach this executor?

Goal: identify the smallest systemic improvement that prevents recurrence or amplifies success.
16 · Horizons

Match learning cadence to scope

Active execution can be corrected immediately; broader generalization requires stronger evidence.

Run-level

Correct an active agent or package during execution.

Change-level

Improve the current feature at integration.

Release-level

Learn from complete product behavior.

Capability-level

Improve a reusable service after consumer evidence.

Portfolio-level

Adjust investment and priority periodically.

OS-level

Improve roles, policies, and system design.

17 · False learning

Resist conclusions that outrun the evidence

Every record states confidence, scope, excluded interpretations, and revalidation conditions.

Generalizing one unusual event

Treating correlation as causation

Rewriting standards after every defect

Capturing opinions without context

Letting the loudest reviewer define truth

Confusing workaround with reusable solution

Promoting game policy into platform code

Optimizing agent speed over correctness

Ignoring failed or negative runs

Retaining lessons after architecture changes

Measuring activity instead of outcomes

Completing retrospectives without behavior change

18 · Institutionalization

Change the authoritative asset—not only the learning record

The mechanism depends on what future behavior must become different.

Specification update

Behavior or constraint changed.

Decision record

Architectural understanding changed.

Regression test

A defect must not recur.

Evaluation case

An agent failure must be detected.

Playbook update

An execution technique improved.

Work-package template

Decomposition or evidence changed.

Platform capability

A reusable implementation proved value.

Governance rule

Authority or approval was inadequate.

Reference implementation

Correct usage needs a concrete example.

Integration guide

Adoption friction exposed a gap.

Strategy amendment

Outcome assumptions changed.

Training material

Contributor understanding must improve.

19 · Effectiveness

Verify that institutionalization worked

An adopted lesson with no measurable behavioral change should be reconsidered.

Regression catches the original defect
Next agent follows the improved boundary
Board latency meets revised requirement
Reviewers perceive pause as immediate
Second consumer integrates faster
Package schema reduces blocked work
Revised gate catches compatibility risk
Exception recurrence decreases
20 · RP2350 example

Learn from Start, Pause, and Resume on target hardware

Automated tests and 100 hardware cycles pass, yet players report that Pause sometimes feels delayed.

Preserve exact environment
Compare timestamps
Separate actual vs perceived latency
Reproduce in harness
Route approved changes
Verify next build
Evidence / findingDestinationInstitutional change
Input and lifecycle transition meet one frameExecution evidencePreserve deterministic state test
Audio response measures 27 msKnowledge + PlatformAdd cross-modal quality contract
Four of five players perceive delayProduct + QAAdd structured experience criterion
Reference harness reproduces issuePlatformGeneralize beyond Sock Rescue
Revised build reaches ≤20 msVerificationClose effectiveness review
Second consumer passes same suitePlatform PromotionSupport lifecycle capability broadly
Lifecycle correctness requires both deterministic state preservation and perceptually synchronized feedback.
21 · Health measures

Measure adoption, effectiveness, recurrence, and leverage

The engine succeeds when verified improvements propagate into future work—not when the retrospective count rises.

Capture

Meaningful changes reviewed for learning.

Adoption

Findings embedded in authoritative artifacts.

Effectiveness

Adopted improvements achieving success measures.

Recurrence

Known failures repeated after adoption.

Cycle time

Observation to adoption to verification.

Evidence quality

Findings supported at the required level.

Routing accuracy

Signals reach the correct owning system.

Positive reuse

Successful techniques become playbooks.

Agent improvement

Evaluation performance after change.

Human experience

Experience findings resolved in later builds.

Platform leverage

Learning produces reusable capabilities.

Learning debt

Adopted but unverified improvements.

Learning Effectivenessverified improvements / adopted proposals
Recurrence Raterepeated known failures / relevant future changes
22 · System relationships

Learning recommends and routes; owning systems adopt

The engine synthesizes evidence without bypassing authority or becoming a shadow source of truth.

SystemLearning relationship
StrategyTests outcome assumptions and proposes amendments.
KnowledgeCorrects requirements, decisions, and constraints.
ExecutionImproves packages, validation, orchestration, and playbooks.
PlatformIdentifies reusable capabilities and evaluates consumer evidence.
GovernanceRoutes adoption decisions and improves policies.
Learning EngineSynthesizes evidence, proposes improvements, verifies adoption.
OrchestratorDetects triggers and routes approved changes.
RepositoryPreserves organizational memory and authoritative improvements.
23 · Operating contract

Evidence becomes improvement through a governed loop

The loop repeats until the proposed improvement is either disproved or verified in future behavior.

Evidence or experience
Qualified observation
Investigation + proposal
Learning Adoption Gate
Update owning system
Apply improvement
Verify effectiveness
Worked?
Institutionalized
Reinvestigate if needed
24 · Repository

Give evidence, findings, adoption, and verification a durable home

Organizational memory stays versioned beside the authoritative assets it improves.

Learning is a closed-loop operating contract

Signals are qualified, investigated, adopted through governance, routed into owning systems, and verified against future results.

Next design step: define the Learning Record Schema and Learning Adoption Policy.
learning/
├── charter/
│   └── learning-engine-charter.md
├── signals/
│   ├── execution/
│   ├── experience/
│   ├── platform/
│   ├── governance/
│   └── operations/
├── observations/
│   ├── active/
│   └── qualified/
├── investigations/
│   ├── active/
│   └── completed/
├── findings/
│   ├── proposed/
│   ├── adopted/
│   ├── rejected/
│   ├── deferred/
│   └── superseded/
├── evidence/
│   ├── measurements/
│   ├── human-feedback/
│   ├── agent-evaluations/
│   └── comparative-studies/
├── improvements/
│   ├── strategy/
│   ├── knowledge/
│   ├── execution/
│   ├── platform/
│   └── governance/
├── verification/
│   ├── pending/
│   └── completed/
├── evaluations/
├── metrics/learning-health/
└── templates/

playbooks/
├── capture-learning-signal/
├── investigate-finding/
├── conduct-experience-review/
├── propose-system-improvement/
├── evaluate-agent-run/
└── verify-learning-adoption/