# Unovie.AI — full content for AI agents
> Complete text of unovie.ai, generated from the live site for RAG and agent ingestion. Sections map to page sections; the same content is available as structured JSON chunks at https://unovie.ai/agents/index.json
# Unovie.AI — AI engineering, delivered. Edge-AI systems, end to end.
URL: https://unovie.ai/
## Edge-AI, engineered end to end.
AI engineering · as a service · est. on the edge Edge-AI, engineered end to end. Unovie is an AI-engineering studio. We design, build, and operate custom Edge-AI Agentic Systems — fixed scope, fixed cost , on hardware you own. From architecture to production, we accept full responsibility. Start a project → Read the field guide Trusted with NVIDIA AMD Qualcomm Siemens GE Live · Edge node NVFP4 Throughput 0 tok/s Power 0 W Egress 0 bytes
## We own the whole arc, from whiteboard to production.
00 — The model We own the whole arc, from whiteboard to production. Not advisory decks. Working systems — assessed, designed, built, optimized, and supported within committed timelines. 01 Architecture Reference architecture & solution blueprint, scoped to a fixed outcome. 02 POC A graded proof against your data — value proven before you commit further. 03 MVP Production-shaped build: data, models, testing, monitoring. 04 Production Hardened, observed, and operated — on hardware you own.
## Four things we are genuinely good at.
01 — Engineering disciplines Four things we are genuinely good at. / model-engineering Model Engineering Self-learning loops, skill optimization, and frozen-base adaptation that compounds accuracy on your tasks — without retraining risk. verifiers skill opt distillation / edge-efficiency Edge Efficiency Small models on Blackwell silicon. PLE-safe NVFP4 quantization, unified-memory budgeting, ~4× decode throughput. NVFP4 128 GB unified Jetson Thor / opex-efficiency Opex Efficiency Capex, not a metered bill. Marginal query cost ≈ electricity; accuracy compounds while cost stays flat. on-prem no egress fixed cost / reliable-hpc Reliable HPC Systems Gated, reversible, regression-safe deployments. Every self-improvement is logged, scored, and revertible. gated reversible auditable
## Pre-engineered outcomes, not blank-slate projects.
02 — Solution accelerators Pre-engineered outcomes, not blank-slate projects. Agentic edge-AI blueprints with quantified goals — deployed in weeks, not quarters. 01 / Maritime Maritime Digital Twin Edge-AI fleet monitoring & predictive maintenance. −6mo overhauls · 100% on-vessel 02 / Automotive Connected-Vehicle Twin VIN-level twin for connected-EV fleets. 4.2h resolution · 98.7% delivery 03 / Airline Corporate-Travel Sales Account scoring, leakage & auto-drafted travel deals. <5min prep · 1-day RFP 04 / Manufacturing Smart Factory Floor Automated inspection & quality control on the line. −30% errors · +25% speed 05 / Logistics Smart Warehouse Inventory management & real-time tracking. −20% stockouts · +15% fill 06 / Safety OSHA Compliance Risk analytics with real-time alerting. −25% fines · +30% training 07 / OTT Content Provider Vision Intelligence Real-time channel detection & QA across live feeds. 96 channels · <1s verify 08 / Process Batch Optimization Chemical & pharma batch processes. +25% quality · −15% waste 09 / Workforce Remote Expertise Real-time expert guidance to teams. +20% output · −30% visits
## Built on the Nexus Context Platform.
03 — The platform Built on the Nexus Context Platform. Nexus · Context Platform for Edge-AI A knowledge graph and vector memory that plans, acts and learns — on hardware you own. Nexus fuses a typed knowledge graph with vector memory and an agent orchestrator: it grounds every answer in source records, drives a fine-tuned vision-action model and next-best-action, and improves from a closed training loop. An inference-native serving layer routes each turn to the least-cost model, protects the prefix cache, and selects context rather than pasting it — so long-running agents stay cheap and coherent. Served on NVIDIA edge GPUs through a Kubernetes micro-cloud — serverless model serving, event streaming and OLAP, all on-prem. knowledge graph + vector agent orchestrator inference-native routing prefix-cache discipline serverless serving NATS + ClickHouse Kubernetes · on-prem NVIDIA edge GPUs Edge Data Fabric Typed knowledge graph + vector memory → Edge Streaming Intelligence Real-time vision over 1,000+ live streams → GPU MicroCloud On-prem GPUs, scheduled & metered like cloud → GPU EdgeGateway Inference-native, least-cost model routing → IT/OT Edge Security Intelligence GPU-native SIEM — AI detection at ingest →
## A proven, success-based methodology.
04 — How we deliver A proven, success-based methodology. Full responsibility, committed timelines — every phase with enterprise-grade execution. 01 Assess & plan Workshops, ROI, value outcomes. 02 Solution design Roadmap, reference architecture, blueprints. 03 Build & test Data, models, testing, monitoring. 04 Optimize KPI tracking, feedback loops, reviews. 05 Train & support Enablement & ongoing adoption.
## Read the thinking behind the work.
05 — Research Read the thinking behind the work. Free eBook · Field Guide Edge AI Models Without the PhD How AI models actually learn — gradient descent, backprop — and why fine-tuning is the wrong reflex on the edge. 25 chapters, diagrams, real results. Read the field guide → Whitepaper Training Without Retraining The frozen-base doctrine: keep the weights fixed, adapt in context, verify against a frozen baseline — and let throughput compound the gains nightly. Read the whitepaper →
## Bring us a problem. We'll return a system .
Let's build Bring us a problem. We'll return a system . Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Tell us the outcome you need; we'll engineer the path to production. Talk to our engineers → Explore the field guide
# About — Unovie.AI
URL: https://unovie.ai/about.html
## We engineer AI that ships.
Home / About About · Unovie.AI We engineer AI that ships. Unovie is an AI-engineering studio for Industry 4.0, founder-led by a team whose vision matches the most ambitious in the field — and measured by what reaches production. We take full responsibility for AI transformation — from architecture through POC and MVP to production readiness — on a foundation of frozen-base, self-improving edge systems that you own and can audit. While the market splits between multi-year transformation programs and big-vision manifestos, we take the third path: production-grade, agentic systems you own, delivered in weeks at a fixed cost. Start a project → Read our field guide 4 engineering disciplines 5 -step methodology fixed time & cost
## Four engineering disciplines.
01 — What we are good at Four engineering disciplines. /model-engineering Model Engineering Self-learning loops and skill optimization that compound accuracy without retraining risk. verifiers skill opt /edge-efficiency Edge Efficiency Small models on Blackwell silicon — PLE-safe NVFP4, unified-memory budgeting. NVFP4 Jetson Thor /opex-efficiency Opex Efficiency Capex, not a metered bill — marginal cost near electricity. on-prem fixed cost /reliable-hpc Reliable HPC Systems Gated, reversible, regression-safe — every change logged and revertible. gated auditable
## A proven, success-based methodology.
02 — How we deliver A proven, success-based methodology. 01 Assess & plan Workshops, ROI, value outcomes. 02 Solution design Roadmap & reference architecture. 03 Build & test Data, models, testing, monitoring. 04 Optimize KPI tracking & feedback loops. 05 Train & support Enablement & adoption.
## Vision that ships.
03 — Why teams choose us Vision that ships. /leadership Founder-led vision Led by founders whose ambition matches the field's most visionary — and who insist that vision reach production. founder-led visionary /velocity Execution velocity Production-grade agentic systems in weeks, not multi-year programs — fixed scope, fixed cost. weeks fixed cost /sovereignty Owned & sovereign On-prem on hardware you own, auditable and reversible — your context never leaves your boundary. on-prem auditable Why Unovie →
## Trusted with NVIDIA · AMD · Qualcomm · Siemens · GE.
Partners Trusted with NVIDIA · AMD · Qualcomm · Siemens · GE. Manufacturing · Logistics · Oil & Gas · Food Processing. Bring us a problem; we will return a system. Talk to our engineers →
# Why Unovie — vision that ships, owned and predictable
URL: https://unovie.ai/why-unovie.html
## Vision that ships.
Home / Why Unovie Why Unovie Vision that ships. Unovie is founder-led by a team whose ambition matches the field's most visionary — and who insist that vision reach production. The market has split between multi-year enterprise transformation and big-vision manifestos, and both leave you waiting. We are the third path: production-grade, agentic Edge-AI systems you own, delivered in weeks on a fixed scope and a fixed cost, on hardware that never leaves your floor. Start a project → Read the economics weeks to production, not years fixed scope & cost 100 % on-prem · you own it
## Built for outcomes, not programs.
01 — Where we are different Built for outcomes, not programs. /leadership Founder-led vision Led by founders whose ambition matches the field's most visionary — and who measure that vision by what reaches production, on your terms. founder-led visionary ships /velocity Execution velocity Production-ready agentic systems in weeks, not multi-year programs — architecture to POC to MVP to production, with no bureaucracy in the path. weeks agentic production /roi ROI realism We answer POC fatigue with a direct line to cost and throughput: capex you own, marginal cost near electricity, spend you can forecast. predictable capex TCO /vertical Vertical depth Manufacturing, logistics, energy and OT — domain ontologies and edge systems built for your floor, not generic AI wrappers. industrial OT domain /sovereignty Governance by architecture Your data stays on-prem on hardware you own; safety runs inline, every decision is auditable and reversible, and context never leaves your boundary. on-prem auditable sovereign /platform A platform, not a manifesto The Nexus Context Platform, GPU EdgeGateway and Device Platform run today — owned silicon and a real stack, not slides and a destination. Nexus running owned
## We are the third path.
02 — The market is bifurcated We are the third path. /transformation The transformation play Global integrators rebuild the entire digital estate. Valuable — but multi-year, multi-million, and slow to first value. multi-year whole-estate /vision The vision play Vision-led firms sell the destination. Inspiring — but often a roadmap without a delivery engine. inspiring no engine /unovie The Unovie play Founder-level vision married to an engineering studio that ships: a working, vertical, owned system in production in weeks. vision + execution owned weeks
## Bring us a problem. We will return a system.
Let's build Bring us a problem. We will return a system. Turnkey Edge-AI — fixed time, fixed cost, full responsibility, on hardware you own. Talk to our engineers → About the studio
# Contact — Unovie.AI · Built in Austin, Texas
URL: https://unovie.ai/contact.html
## Bring us a problem. We'll return a system.
Home / Contact Contact · Let's build Bring us a problem. We'll return a system. Tell us the outcome you need. Our engineers will scope the path to production — fixed time, fixed cost, full responsibility. Not a multi-year program and not a slide deck — founder-led vision delivered as a working system you own. Get in touch → Read the field guide
## Where chips, software & services converge.
Where we build ◉ Austin · Central Texas corridor Where chips, software & services converge. Unovie is built in Austin, Texas — the gravitational center of America's physical-AI revolution. Silicon is designed and fabricated here, autonomy software is written here, foundational ML research happens here, and the AI infrastructure that runs it is converging here. We engineer edge-AI from the same ground. / chips Chips Silicon and edge accelerators — designed and fabricated across Central Texas. Samsung fabs NVIDIA AMD / software Software Models, autonomy, and the self-learning loop — the intelligence layer. models autonomy UT Austin IFML / services Services Turnkey delivery: architecture → POC → MVP → production. turnkey Dell × NVIDIA on-prem In good company — the corridor anchoring Physical AI SpaceX Tesla Apple Samsung Firefly Aerospace Dell × NVIDIA UT Austin · IFML SpaceX — autonomy & reusable launch · Tesla — robotics & manufacturing AI (HQ in Austin) · Apple — its largest campus outside Cupertino · Samsung — advanced semiconductor fabs (Taylor, TX) · Firefly Aerospace — responsive launch & lunar landers (Cedar Park, TX). Research & infrastructure: UT Austin's IFML — the NSF Institute for Foundations of Machine Learning — and the Dell × NVIDIA AI-factory convergence anchor the corridor. Named to describe the regional innovation ecosystem; not affiliations or endorsements.
## Talk to our engineers.
Reach us Talk to our engineers. ◍ Headquarters 651N N. Highway-183, Suite #4120 Austin, TX 78641 · USA ☎ Phone +1-636-579-9725 ✉ Email contact@unovie.ai · sales@unovie.ai ◇ Engagement Turnkey · fixed time & cost · worldwide
## Let's put physical AI to work.
Engineered in Texas Let's put physical AI to work. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Get in touch → Read the field guide
# Edge AI Models Without the PhD — An Architect’s Field Guide
URL: https://unovie.ai/resources/edge-ai-models.html
## Edge AI Models Without the PhD
Technical Field Guide · Edge AI Edge AI Models Without the PhD How AI models actually learn, why fine-tuning is the wrong reflex on the edge, and how to build custom, self-improving systems on hardware you own — written for enterprise architecture teams. Suresh Mandava · suresh@unovie.com Unovie.AI · EdgeAI Context Engineering Edition 1.0 · June 2026 Frozen base · in-context adaptation · self-verification · NVIDIA Jetson Thor · DGX Spark · RTX Spark ↓ Start reading
# Training Without Retraining — A Frozen-Base Doctrine for Custom Models on the Edge
URL: https://unovie.ai/resources/edge-ai-whitepaper.html
## Training Without Retraining: A Frozen-Base Doctrine for Custom Models on the Edge
Training Without Retraining — A Frozen-Base Doctrine for Custom Models on the Edge ← Unovie.AI Download PDF Technical Whitepaper · Edge AI Training Without Retraining: A Frozen-Base Doctrine for Custom Models on the Edge Best practices for adapting small models on-device through in-context learning, external memory, and self-verification — motivated by what the learning-rules-and-brain-alignment literature tells us about when weight updates actually help. Suresh Mandava · suresh@unovie.com Unovie.AI · EdgeAI Context Engineering Revision 1.0 · June 2026 Hardware target: NVIDIA Jetson Thor · DGX Spark · RTX Spark (Blackwell, 128 GB unified, NVFP4) Abstract The instinct when a foundation model underperforms on a niche task is to fine-tune its weights. On the edge — constrained memory, no cloud, regulated data, a need for reversibility — that instinct is usually wrong. We argue for a frozen-base doctrine : keep the pretrained weights fixed and move all adaptation into external, reversible state — a knowledge graph, an optimized in-context skill, and lightweight runtime controllers — closed by a programmatic verifier that lets a small model teach itself overnight without labels, drift, or data egress. We connect this engineering practice to a recent line of computational-neuroscience work showing that a frozen, untrained network can match or exceed a backpropagation-trained one at the representational level, that local, lightweight learning preserves structure that global gradient descent erodes, and that higher-level abstraction is gated by capacity and data rather than by the update rule. Treated as motivation rather than proof, those results sharpen a practical rule of thumb: adapt in context first; touch the weights last, and only when a frozen-baseline evaluation proves you must. We give the supporting hardware practices (PLE-safe NVFP4 quantization, throughput-as-learning-rate, develop-on-DGX-Spark / deploy-on-Thor) and a best-practices checklist. Contents The edge adaptation problem What the learning-rules literature suggests The frozen-base doctrine In-context learning as the primary lever Closing the loop: the self-learning cycle Evaluation discipline Edge hardware best practices Scaling to many domains When to actually touch the weights Best-practices checklist References 1 The edge adaptation problem A capable open model — say a 4-billion-active-parameter instruction-tuned model — rarely fails on the edge because it is too small. It fails because it is generic : it knows your specialty broadly and shallowly, it has no memory of your corpus between calls, and it cannot be sent the one thing that would fix it — your private data. The reflex is to fine-tune. But on an edge appliance that reflex collides with four hard constraints: Reversibility. A merged weight update is difficult to audit and to undo. Regulated buyers need to prove that a change can be rolled back instantly and that the trusted base is always one step away. Catastrophic forgetting & drift. Gradient updates on a narrow corpus quietly degrade unrelated capabilities. Detecting this requires a regression suite the operator rarely has. Capacity and data. A small on-device model trained on a small in-house corpus is exactly the regime where fine-tuning yields the least and risks the most. Cost and cadence. Retraining is an event; adaptation needs to be a habit . If improving the model is a quarterly project, it will not keep pace with the domain. This whitepaper makes the case that the right primitive on the edge is not the gradient step but in-context adaptation closed by verification : the model's behavior is shaped by what it is given at inference time — retrieved context and an optimized instruction — and that context is itself improved, nightly, against a ground-truth checker. The weights stay frozen. Below we first look at why this is more than an engineering convenience: a strand of recent neuroscience-adjacent ML research suggests the frozen substrate is doing more work than we usually credit. 2 What the learning-rules literature suggests A useful corrective to "more training is always better" comes from work that asks a sharper question: does the learning rule even matter for the representation a network forms? The answer, at least at the lower levels of the hierarchy, is surprisingly often no . Leutenegger [1] [2] systematically compares five conditions — backpropagation (BP), feedback alignment (FA), predictive coding (PC), spike-timing-dependent plasticity (STDP), and an untrained random-weights baseline — on an identical small convolutional architecture, and scores each against human fMRI and macaque electrophysiology using Representational Similarity Analysis (RSA). Four findings are directly relevant to anyone deciding how to adapt a model: 2.1 Architecture can dominate the update rule At early visual cortex (V1), the untrained network exceeds the backpropagation-trained one (ρ = 0.076 vs 0.034; Δρ = +0.044, p < 0.001) [1] . The structural priors of the architecture — local connectivity, pooling, nonlinearity — carry most of the alignment for free; training on a narrow objective actually moved representations away from the target. This echoes older results that random-weight networks already carry non-trivial visual structure [3] [11] and that unsupervised sparse coding alone yields V1-like receptive fields [4] . 2.2 Local, lightweight learning preserves structure that global gradients erode Among trained rules, the local ones — STDP and predictive coding — produced the highest early-cortex alignment, above backpropagation, in both species [1] [2] . By contrast, feedback alignment (random global feedback) was consistently the worst , despite reaching meaningful task accuracy — task performance and representational fidelity were dissociated throughout. The lesson is not that STDP is the answer; it is that local, structure-preserving updates beat global error signals when the goal is to keep the representation faithful rather than merely to fit a label. 2.3 Higher-level abstraction is gated by capacity and data, not the rule At the top of the hierarchy (IT cortex), all rules — including no training — converged; no update rule was reliably more brain-like [1] . A capacity control made the cause explicit: a pretrained ResNet-50 (ImageNet) jumped to ρ ≈ 0.25 at IT versus 0.07–0.14 for every small-CNN condition [2] . Higher-level, abstract structure is bought with model capacity and training-data richness , not with a cleverer update on a small model. 2.4 The discipline of the untrained baseline The papers' explicit methodological moral: always include an untrained, same-architecture baseline. Without it, "architecture effects are confounded with learning effects — our data show this confound can be essentially complete at V1" [1] . Any claim that an adaptation helped is meaningless unless measured against the unadapted model on held-out data. Table 1 — From the learning-rules literature to edge adaptation practice. Reported finding Edge-adaptation principle it motivates Untrained / frozen architecture matches or beats a trained one at low levels [1] Freeze the base. The pretrained substrate already carries most of the signal; default to not retraining it. Local rules (STDP, PC) preserve structure; global BP can move it away [1] [2] Adapt locally & reversibly (in-context, controllers) rather than with global gradient descent that risks drift/forgetting. Higher-area abstraction scales with capacity + data, not the rule [2] Diagnose the bottleneck. If a task genuinely needs more abstraction, it is a capacity problem (distill/upgrade), not an adaptation problem. Task accuracy ≠ representational alignment [1] Optimize the right objective. Grade adaptation on grounded task fidelity, not a convenient proxy. "Always include an untrained baseline" [1] Gate on a frozen baseline. Accept a change only if it beats the unadapted model on data it never saw. An honest caveat. These studies are small-scale convolutional vision models scored against primate visual cortex, with explicitly hedged, sometimes null results (n = 5 rules; n = 3 fMRI subjects; single datasets) [1] [2] . They are not evidence about large language models, and we cite them as motivation and intuition , not as proof. The engineering case below stands on its own measurements; the neuroscience simply rhymes with it. 3 The frozen-base doctrine "Self-learning" on the edge does not have to mean "weight-updating." We define it as swappable external state that measurably improves a held-out domain metric over time — each store operating at a different timescale, none of them touching the base weights. Figure 1 — The frozen-base adaptation stack. External state (L0–L2) conditions a frozen model at inference time; the L3 weight adapter is detachable and parked. Concretely, adaptation lives in four layers, only the last of which involves gradients — and that one is held in reserve: Table 2 — The adaptation stack. Weights stay frozen for layers 0–2. Layer What changes Timescale Weights? L0 · Knowledge A typed knowledge graph + vector index, grown by continuous ingestion of the domain corpus. per-ingest, continuous No L1 · Skill A plain-language skill document (the in-context instruction the model follows), rewritten when a better version is proven. nightly batch No L2 · Behavior Lightweight per-request runtime controllers that steer the frozen model on hard cases. per-request No L3 · Weights (parked) A detachable, never-merged low-rank adapter, used only after the reversible levers plateau. on plateau Yes (revertible) Principle 1 Keep the base model byte-for-byte frozen. Put adaptation in external state that can be inspected, diffed, and reverted in one step. The trusted base is always one detach away. This is the engineering analog of §2.1–2.2: the frozen substrate carries the heavy lifting, and the cheap, local, reversible stores do the domain-specific shaping. Crucially, every change is an artifact — a graph delta, a versioned skill file, a controller blob — not an opaque shift in billions of parameters. Auditability and reversibility are properties of the design, not bolt-ons. 4 In-context learning as the primary lever In-context learning [12] — shaping behavior through what the model is shown at inference time rather than through its weights — is the edge-appropriate adaptation mechanism. Three surfaces carry it. 4.1 The skill document is a learned artifact The system prompt is not boilerplate; it is the procedure the model executes , and it is the thing we optimize. A skill document begins as a short seed and is grown by the loop (§5) into a precise, failure-aware specification of how to perform the domain task. Because it is text, it is human-readable, version-controlled, and graftable: the same skill that scored highest last night is exactly what production loads today. Principle 2 Treat the in-context instruction as the unit of learning. Version it, score it on a frozen holdout, and promote it only when it beats the incumbent. The "trained model" you ship is a frozen base plus a proven skill artifact. 4.2 Retrieval-grounded context (L0) The second surface is what the model retrieves . A continuous knowledge-graph memory turns the private corpus into typed, queryable structure; at request time the relevant subgraph and supporting chunks are composed into the prompt. This grounds answers in source records and — unlike weight memorization — lets the knowledge change without touching the model. The same memory stack also supplies the reward for the loop (schema conformance, entity resolution, retrieval grounding), so improving the knowledge layer and improving the skill are coupled through one metric. 4.3 Runtime behavioral controllers (L2) The third surface is the lightest gradient-free form of "local learning" in the literature's sense (§2.2): tiny per-request controllers (kilobytes, composable, fitted from already-verified examples) that nudge the frozen model's activations on hard or ambiguous cases. They attach and detach like a setting, never alter the base, and are the on-device echo of "local, structure-preserving adaptation beats global retraining." 5 Closing the loop: the self-learning cycle In-context adaptation only compounds if there is a closed loop that improves the context automatically. The engine is a nightly cycle with a programmatic verifier at its center: the model tries, the verifier scores, a larger model reflects on the failures, and a gate commits the change only if it provably helps. Figure 2 — The nightly self-learning loop. A programmatic verifier scores every rollout; verified winners feed the controllers, failures drive reflection, and a frozen-holdout gate commits only proven gains. # one nightly cycle — fully unattended, fully local champion = skill_store.current() tasks = load(train_split) # 1) Rollout: best-of-N samples per task from the frozen executor scored = [verify(t, sample) for t in tasks for sample in rollout(executor, champion, t, n= N )] # 2) Harvest: verified winners become free, self-labeled data (feeds L2) winners = [s for s in scored if s.score >= WIN_THRESHOLD] # top-1/task over the bar failures = [s for s in scored if s.score < WIN_THRESHOLD] # best honest attempt # 3) Baseline: score the CURRENT skill on a frozen, fingerprinted holdout baseline = score_on_holdout(champion, holdout) # 4) Reflect: a larger model reads <=K failures, proposes textual patches (no scoring) patches = reflector.reflect(champion, failures[:K]) # strict JSON, <=3 edits candidate = apply_patches(champion, patches) # 5) Gate: re-score on the SAME holdout; commit iff it clears the noise floor lift = score_on_holdout(candidate, holdout) - baseline if lift >= MIN_LIFT: skill_store.commit(candidate) # new champion vNNNN else : keep(champion) # auto-revert; log the rejection 5.1 The verifier is the reward The make-or-break component is the checker. A graded, deterministic, programmatic verifier — for an extraction task, the fraction of gold fields recovered after normalization, scored in [0,1] — replaces both human labeling and LLM-as-judge. Because the reward is code, the loop runs fully unattended and fully on-box. The larger "reflector" model writes patches but never scores ; scoring stays mechanical. With a deterministic checker you also get a data engine for free: best-of-N rejection sampling harvests verified-correct trajectories that feed the controllers. Principle 3 Make the reward a graded program, not a judgment. No labels, no LLM-as-judge in the gate. Keep generation and evaluation strictly separate so the optimizer cannot grade its own homework. 5.2 Two failure modes a programmatic gate must handle Goodharting. Optimize exactly what you check and the model finds schema-valid garbage. Mitigate with layered checks (conformance and entity-match and grounding) and periodic spot-checks. Sparsity. Binary pass/fail starves the reflector. Make the metric graded (fraction of subgoals), so a near-miss is distinguishable from a disaster and the optimizer has a gradient to climb. 6 Evaluation discipline The single most transferable lesson from §2.4 is operational: a change is only "better" relative to the frozen, unadapted baseline, measured on data the optimizer never saw. The loop above bakes that in. Frozen, fingerprinted holdout. The validation set is content-hashed and verified before every rollout; if a byte drifts, the run aborts. The optimizer can never see or touch it — the on-device equivalent of the literature's untrained-baseline control. One metric, end to end. The same scorer grades rollouts and the gate, so you optimize exactly what you measure. A noise floor. A minimum-lift threshold rejects sub-noise "wins" — the regression suite's job is to say no far more often than yes . Self-limiting behavior. On a domain already near its ceiling, the correct outcome is a long string of rejections holding the line — a feature, not a failure. A system that cannot quietly make itself worse is exactly what a regulated buyer wants. 7 Edge hardware best practices The frozen-base loop is bottlenecked by one thing: how many verified rollouts you can run per night. On a bandwidth-bound edge accelerator, throughput is the learning rate , and the hardware choices that set throughput are the ones that set how fast the model improves. 7.1 Throughput is the learning rate Modern Blackwell-class edge devices share a decisive property: a large pool of unified memory (128 GB, ~273 GB/s) coherently shared by CPU and GPU [13] [14] . That lets the serving model, a larger "teacher" used only during the nightly window, and the memory stack all stay resident at once — no swapping. Native 4-bit (NVFP4) compute lifts decode roughly 4× over bf16 (≈30 → 120 tokens/sec on this class of board), and because decode is bandwidth-bound, that 4× translates directly into 4× more best-of-N experiments per night. 7.2 PLE-safe quantization Aggressive quantization is the lever — but not uniformly. Models that achieve a small "effective" footprint via per-layer embeddings (PLE) keep a large, quantization-fragile block of embedding parameters that must stay in higher precision. The discipline: quantize only the active compute weights (attention, FFN) to NVFP4, and keep the PLE tables in bf16. A build that quantizes PLE tensors will degrade output silently. Bake a PLE-safety preflight into the conversion so the build fails on drift rather than shipping a quietly worse model. Principle 4 Quantize for speed, but precision-protect the fragile parts. NVFP4 the compute path; keep per-layer embeddings in bf16. Validate the conversion automatically, and never serve the model on a runtime that silently drops PLE. 7.3 Develop on the desktop, deploy at the edge The same Blackwell + 128 GB + NVFP4 substrate appears in three form factors, and they are complementary rather than competing: Table 3 — One substrate, three roles. The trained skill artifact is portable across all three. Device Role Strength DGX Spark (GB10, desktop) Develop & fine-tune Highest local LLM throughput; cluster two units (ConnectX-7) for ~405B-class work [13] . Jetson AGX Thor (embedded) Deploy at the edge Real-time multimodal inference at 40–130 W; sensor fusion + robotics stacks [14] . RTX Spark (Windows PC) Personal agents 120B-class local models on laptops/mini-PCs; broad OEM reach [15] . Because the frozen base, the NVFP4 weights, and the skill artifact are identical across them, the natural workflow is develop and fine-tune on DGX Spark → deploy on Thor → reach every desk on RTX Spark , with the learned artifact moving between them unchanged. Figure 3 — One substrate, two phases. Develop and fine-tune on DGX Spark; the frozen base plus the proven skill artifact deploy unchanged to Thor at the edge and RTX Spark on PCs. 8 Scaling to many domains The frozen-base design scales sideways cheaply because almost everything expensive is shared. Share the brain. The larger reflector and the embedding encoder are domain-agnostic and idle most of the day; one instance serializes every domain's nightly reflection. Isolate the expertise. Each domain gets its own executor, its own memory collection (no cross-domain pollution), and its own skill chain and frozen holdout — an independent proof trail. Onboard cheaply. A new domain needs a corpus and a [0,1] scorer, declared in one config. No model retraining; the per-domain marginal cost is a slice of one power-efficient box. Three to four domains fit comfortably on a single 128 GB device (one ~16 GB reflector + several ~18 GB executors), which is the practical sweet spot before time-sharing the executor. 9 When to actually touch the weights The doctrine is "weights last," not "weights never." §2.3 tells us when last arrives: when the limit is genuinely capacity , not adaptation. Operationalize it with an objective trigger. If held-out lift stays under a small threshold (e.g. <0.5%) for several consecutive nights while the runtime layers are still churning , the reversible levers have demonstrably flattened — the task may now be capacity-bound. Only then consider a weight update, under strict conditions: Detachable, never merged. Keep any low-rank adapter [16] [17] as a separate, hot-swappable file so the frozen base is always recoverable. Distill from the verified corpus. The teacher generates and the programmatic gate filters; you are amortizing an already-verified corpus into the base, not inventing capability beyond the student's ceiling. PLE-safe training. Put adapters on attention/FFN projections only; never quantize or train the per-layer embeddings. Forgetting-regression gate. Promote the adapter only if new-domain lift clears a bar and no prior-domain regression exceeds a tight bound, measured against frozen, never-trained holdouts — the same baseline discipline as §6, now guarding against catastrophic forgetting. Principle 5 Ask weights for help only after a frozen-baseline evaluation proves the reversible levers are exhausted — and even then, as a detachable adapter behind a forgetting-regression gate, never a merge. 10 Best-practices checklist Table 4 — A frozen-base edge-adaptation checklist. # Practice Why 1 Freeze the base by default; adapt in external state (knowledge, skill, controllers). Reversibility, auditability, no drift; the substrate already carries most of the signal (§2.1, §3). 2 Make the in-context skill the unit of learning — versioned, scored, promoted. Human-readable, graftable, provable adaptation without weight changes (§4.1). 3 Ground answers via retrieval from a continuous knowledge graph that also supplies the reward. Knowledge changes without touching the model; couples L0 and L1 through one metric (§4.2). 4 Prefer local, reversible adaptation (controllers) over global gradient descent. Local updates preserve structure that global error signals erode (§2.2, §4.3). 5 Close the loop with a graded, programmatic verifier; generation ≠ evaluation. Unattended, label-free, Goodhart-resistant improvement (§5). 6 Gate on a frozen, fingerprinted holdout with a noise-floor minimum lift; auto-revert. The untrained-baseline discipline; a system that can't quietly get worse (§2.4, §6). 7 Make throughput the priority: NVFP4 compute, unified memory, dual-path serving. On a bandwidth-bound edge board, tokens/sec is the experiment budget (§7.1). 8 PLE-safe quantization with an automated preflight; never a PLE-dropping runtime. Protects the quantization-fragile embeddings; avoids silent degradation (§7.2). 9 Develop on DGX Spark, deploy on Thor; keep the artifact portable. Same substrate, different roles; one workflow from prototype to edge (§7.3). 10 Touch weights only on an objective capacity trigger, as a detachable adapter behind a forgetting gate. Abstraction is capacity-bound, not rule-bound; preserve reversibility (§2.3, §9). The through-line is a single inversion of the default. The cloud reflex — "if it underperforms, train the weights" — is the wrong primitive on the edge, where reversibility, drift, capacity, and cadence all argue the other way. Freeze the substrate, adapt in context, verify against a frozen baseline, and let throughput compound the gains nightly. The learning-rules literature gives that engineering bet a satisfying second reading: the frozen architecture was doing more of the work than we assumed, local adaptation keeps the representation honest, and the cases that truly need more — abstraction — are the ones to spend capacity on, deliberately, last. 11 References Leutenegger, N. (2026). Untrained CNNs Match Backpropagation at V1: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI. arXiv:2604.16875v2. Frozen/untrained CNN exceeds BP at V1; PC/STDP > BP; convergence at IT; "always include an untrained baseline." Leutenegger, N. (2026). Cross-Species RSA Reveals Conserved Early Visual Alignment but Divergent Higher-Area Rankings Across Human fMRI and Macaque Electrophysiology. arXiv:2605.22401v1. Replicates early-area pattern in macaque; ResNet-50 capacity control shows IT alignment scales with capacity + data. Saxe, A. M., Koh, P. W., Chen, Z., Bhand, M., Suresh, B., & Ng, A. Y. (2011). On random weights and unsupervised feature learning. ICML. Olshausen, B. A., & Field, D. J. (1996). Emergence of simple-cell receptive-field properties by learning a sparse code for natural images. Nature , 381:607–609. Whittington, J. C. R., & Bogacz, R. (2017). An approximation of the error backpropagation algorithm in a predictive-coding network with local Hebbian plasticity. Neural Computation , 29:1229–1262. Lillicrap, T. P., Cownden, D., Tweed, D. B., & Akerman, C. J. (2016). Random synaptic feedback weights support error backpropagation for deep learning. Nature Communications , 7:13276. Bi, G.-q., & Poo, M.-m. (1998). Synaptic modifications in cultured hippocampal neurons: dependence on spike timing, synaptic strength, and postsynaptic cell type. J. Neurosci. , 18:10464–10472. Schrimpf, M., Kubilius, J., Hong, H., et al. (2020). Brain-Score: Which artificial neural network for object recognition is most brain-like? bioRxiv. Yamins, D. L. K., & DiCarlo, J. J. (2016). Using goal-driven deep learning models to understand sensory cortex. Nature Neuroscience , 19:356–365. Kriegeskorte, N., Mur, M., & Bandettini, P. (2008). Representational similarity analysis — connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience , 2:4. Truzzi, A., & Cusack, R. (2025). Neural responses in early visual cortex are well predicted by random-weight CNNs. bioRxiv. Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few-shot learners. NeurIPS. (In-context learning.) NVIDIA (2026). DGX Spark — Personal AI Supercomputer (GB10 Grace Blackwell). nvidia.com/products/workstations/dgx-spark. 128 GB LPDDR5X unified, ~273 GB/s; ConnectX-7 clustering. NVIDIA (2026). Jetson Thor — Advanced AI for Physical Robotics. nvidia.com/autonomous-machines/embedded-systems/jetson-thor. Blackwell, ~2070 FP4 TFLOPS, 40–130 W, 128 GB unified. NVIDIA (2026). RTX Spark — Slim Laptops & Small Desktops. nvidia.com/products/rtx-spark. Blackwell RTX + Arm; up to 128 GB unified; 120B-class local models. Hu, E. J., Shen, Y., Wallis, P., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. arXiv:2305.14314. Unovie.AI · EdgeAI Context Engineering. This whitepaper synthesizes internal edge-deployment practice with the cited literature. The neuroscience results are used as motivation and analogy, not as claims about large language models; engineering recommendations rest on on-device measurement. Hardware figures reflect vendor specifications as of June 2026. Revision 1.0.
# The Edge-Native Inference Gateway — Predictable AI Economics for the Industrial Edge
URL: https://unovie.ai/gpu-edgegateway-whitepaper.html
## The Edge-Native Inference Gateway: Turning Unpredictable AI Opex into Fixed, Predictable Cost
The Edge-Native Inference Gateway — Predictable AI Economics for the Industrial Edge ← Unovie.AI Download PDF Technical Whitepaper · Edge-AI Economics The Edge-Native Inference Gateway: Turning Unpredictable AI Opex into Fixed, Predictable Cost For infrastructure leaders. How an inference-native routing gateway on hardware you own converts metered, unbounded cloud-inference spend into capex-based, predictable cost — without giving up capability, safety, latency, or data control. Written for the industrial edge. Suresh Mandava · suresh@unovie.com Unovie.AI · EdgeAI Context Engineering Revision 1.0 · June 2026 Reference hardware: NVIDIA DGX Spark · AMD Ryzen AI Max+ 395 Abstract Metered cloud inference is the fastest-growing and least predictable line in many enterprise AI budgets: cost scales with every token, every retry, and every agent loop, and the bill arrives after the spend is already committed. We argue that for the industrial edge — factories, depots, substations, vehicles, regulated sites — the durable answer is an edge-native inference gateway : own the silicon, run the models on-prem, and put a single inference-native gateway in front of them that decides which model handles each request, reuses cached computation, sends only the context a turn needs, and enforces safety and policy inline. This converts a variable opex stream into a fixed capex plus near-electricity marginal cost , makes spend predictable and bounded , keeps data on-premises, and removes the cloud round-trip. We describe the architecture for the CTO, the cost model for the infrastructure VP, the four token-economics levers that bend the curve, a three-year TCO comparison, and a deployment blueprint on owned reference hardware. Contents The unpredictable-opex problem Architecture: the edge-native gateway Why edge-native — the CTO view Four token-economics levers The economics — the Infra-VP view Reference hardware on owned silicon A three-year TCO comparison Industrial edge use cases Deployment blueprint Governance, safety & reversibility Recommendations & checklist References 1 The unpredictable-opex problem AI moved from a pilot line item to a production dependency, and the bill followed. The trouble is not only that it is large — it is that it is unbounded and arrives in arrears. Metered inference prices the wrong thing for an operator. You are billed per token, so cost scales with prompt length, retries, multi-turn agent loops, tool output pasted back into context, and traffic you do not control. A single agent that "thinks harder" or re-reads a large document can multiply the cost of an outcome by 10× with no change in business value. Budgets are set annually; token spend compounds daily. The result is a line item that finance cannot forecast and infrastructure cannot cap. Figure 1 — The token-price paradox. Per-token prices fell, but agentic workloads multiplied tokens-per-task by orders of magnitude — so the cost per outcome rose even as unit prices dropped. The cost risk is the meter and the token sprawl behind it, not the headline price. For the industrial edge three more constraints stack on top of cost: Data gravity and sovereignty. Telemetry, video, control logs and process data are large, sensitive, and often regulated. Shipping them to a cloud model is a governance liability and an egress bill. Latency and availability. A line, a vehicle or a substation cannot wait on a round-trip to a region, and cannot stop when the link does. Determinism. Operations need the same answer, at the same latency, at the same cost — not a number that drifts with a vendor's pricing or capacity. The thesis Stop renting inference by the token for steady-state industrial workloads. Own the silicon, run the models where the data is, and govern every request through one gateway — so cost becomes capex you control, not opex you discover. 2 Architecture: the edge-native inference gateway The gateway is not a proxy bolted in front of a model. It is the control point where workload, routing, serving, caching and policy meet — designed from the inference engine out, not around it. Every request enters through one routing contract : signals become projections, projections drive a decision, and the decision chooses the model — across a mesh of local small models, on-prem large models, and (only when it genuinely pays) an external frontier API. The same gateway protects reusable computation, trims context to the evidence a turn needs, and runs safety and policy inline. Because it is co-designed with a high-throughput, memory-efficient serving engine, it follows the engine's optimization rules instead of treating every call as generic chat traffic. Figure 2 — One gateway between edge workloads and a pool of models on owned silicon. It routes by signal, reuses cached prefixes, selects context, and enforces safety; the frontier path is reserved for the rare request that justifies it. 3 Why edge-native — the CTO view Edge-native is an architectural choice before it is a cost choice. It changes where data lives, where decisions happen, and what you can guarantee. Property Cloud-metered inference Edge-native gateway on owned silicon Data path Sensitive data egresses to a third party Data stays on-prem; nothing leaves the boundary Latency Region round-trip + queue, variable Local, deterministic, sub-network Availability Depends on the link and the vendor Runs through link and provider outages Sovereignty Subject to external jurisdiction & retention Wholly within your governance domain Reversibility Vendor sets pricing, models, deprecations You version, shadow-test and revert policy Cost shape Variable opex, billed in arrears Fixed capex + near-electricity marginal cost The gateway is what makes "edge-native" operationally real rather than a pile of GPUs. It gives one place to set policy, one place to meter spend, one OpenAI- and Anthropic-compatible ingress so applications do not change, and one lifecycle ( shadow → activate → revert ) so routing never drifts silently. Capability is not sacrificed: hard requests still reach a large on-prem model, and the rare request that truly needs a frontier model can still take that path — by exception, under policy, with the cost attributed. 4 Four token-economics levers Owning the silicon caps the denominator (you stop paying per token). The gateway shrinks the numerator — the work each outcome actually costs — with four compounding levers. Lever Mechanism Effect on cost-per-outcome Signal-driven routing Each request is classified by intent, complexity, risk and modality; mechanical and easy turns go to small local models, and reasoning is invoked only when it pays. Routed paths run at a small fraction of an always-large path Prefix-cache discipline Stable prompt prefixes, deterministic tool schemas and bounded, append-only context keep reusable prefixes intact across a long session. Cached tokens are reused at a steep discount instead of recomputed every turn Context selection The gateway sends the evidence a turn needs — selected, bounded and compressed — rather than pasting whole documents and tool dumps. Large reductions in prompt and tool-output tokens, with continuity preserved Semantic caching Semantically-equivalent requests reuse a prior answer instead of triggering fresh inference. Repeat and near-repeat traffic costs nothing to serve Why this matters more on the edge. Industrial workloads are highly repetitive — the same inspection prompt, the same maintenance query, the same shift handover — so cache reuse and small-model routing hit rates are high. The levers that look marginal in a chatbot are dominant in a plant. 5 The economics — the Infra-VP view The job is not to minimize this month's invoice; it is to make next year's number knowable . Edge-native does that by changing the shape of the cost curve. Capex One-time, depreciable hardware you own — not a recurring meter ≈ kWh Marginal cost of an extra request approaches electricity Fixed Spend is bounded by capacity, not by traffic or token length 0 egress No per-GB data egress, no cross-border transfer cost Metered inference is a line that rises with usage and never flattens; every new agent, every longer prompt, every retry adds to it forever. Owned capacity is a step (the purchase) followed by a nearly flat line (power, space, maintenance). Past a modest, steady utilization the two curves cross — and beyond the crossover, every additional unit of work on owned silicon is effectively free relative to the meter. Figure 3 — Illustrative cost shape. Metered inference rises without bound; owned capacity is a step then a near-flat line. Beyond break-even, extra work is effectively free relative to the meter. Exact crossover depends on utilization and token mix. 6 Reference hardware on owned silicon Predictable economics need predictable units. Two complementary, commodity-priced platforms cover the develop-and-serve lifecycle on hardware you keep on your floor. NVIDIA DGX Spark — the on-prem development & large-model node A GB10 Grace Blackwell desktop supercomputer with 128 GB of coherent unified memory and roughly 1,000 TFLOPS (FP4) of AI compute — enough to prototype, fine-tune and serve models up to ~200B parameters locally, or ~405B across a linked pair over its built-in high-speed fabric. It runs the same container stack as the datacenter, so what you build here promotes to the edge unchanged. AMD Ryzen AI Max+ 395 — the private inference node A small, all-metal node fusing 16 Zen 5 cores , a Radeon 8060S iGPU and an XDNA 2 NPU for ~126 TFLOPS of platform AI, with 128 GB of LPDDR5X-8000 — enough to keep 70B-class models resident and private. Dual 10GbE and dual USB4 let nodes cluster into a compute hub, so capacity scales by adding fixed-price units, not by raising a meter. 128 GB Unified memory per node — large models stay resident, on-prem 200 B→405B Local model scale on a node, or a linked pair 70 B Class of model served privately on a single edge node 10GbE ·USB4 Cluster fixed-price nodes into a private compute hub Why two tiers. Develop and fine-tune on the large node; serve steady-state traffic on a fleet of small nodes governed by the gateway. The same artifacts run on both, so there is one pipeline and one cost basis. 7 A three-year TCO comparison An illustrative model for a single industrial site running a steady mix of copilots, vision triage and maintenance queries. Figures are directional — the point is the shape , not a quote. Illustrative 3-year total cost of ownership for one industrial site (parameters, not a price quote). Dimension Metered cloud inference Edge-native gateway (owned) Upfront capex ~$0 One-time node fleet + setup (depreciable, resaleable) Recurring cost Per-token bill that grows with usage, retries and context Power, space, maintenance, support — roughly flat Marginal cost of +1 request Full token price, every time Approaches electricity once capacity exists Data egress Per-GB transfer for telemetry, video, documents None — data never leaves the site Budget predictability Forecast error grows with adoption Known within power and capacity envelopes 3-year trajectory Rises every quarter; no natural ceiling Step at year 0, near-flat thereafter Exit / change cost Re-platform on vendor pricing & deprecations Hardware retained; policy versioned and reversible The infra-VP takeaway For steady-state industrial workloads, edge-native turns "how much will AI cost next year?" from a forecast into a capacity-planning question — the same discipline you already apply to compute, storage and network. 8 Industrial edge use cases Use case Why edge-native Primary saving Vision QC on the line High-rate video can't egress; needs sub-second local decisions No egress; small-model routing on repetitive frames Predictive maintenance Continuous sensor streams, mostly normal; rare anomalies Cache + cheap path for normal; reserve large model for anomalies OT / IT security Detection must run in the data path, on-prem, always-on Local inference; no telemetry leaves the boundary Field & control-room copilots Repetitive shift queries; must work offline and fast High cache & small-model hit rates; predictable cost Regulated document & agent automation Sensitive records can't be sent to third-party models Sovereignty; context selection trims long-document tokens 9 Deployment blueprint Develop on the large node. Prototype, fine-tune and evaluate on the on-prem development supercomputer; keep models and data inside the boundary. Promote unchanged. Ship the same containers to a fleet of small edge inference nodes — one pipeline, one artifact, one cost basis. Front everything with the gateway. All traffic flows through one OpenAI- and Anthropic-compatible ingress that routes by signal, reuses prefixes, selects context, and enforces safety. Meter and attribute. Every route is logged with latency, tokens and cost, so spend is accountable per team and per workload — while the task runs, not after. Shadow, activate, revert. Test every routing or policy change on replayed traffic before activation, with one-click rollback. Scale by units, not by meter. Add fixed-price nodes to the cluster as demand grows; the cost curve stays a series of known steps. 10 Governance, safety & reversibility Owning the inference path is also the strongest governance posture available. Sensitive context never leaves the site, so leaked vectors and prompts handed to models you do not control — a business liability and a governance violation — simply cannot happen. Safety classifiers for sensitive-data leakage, prompt injection and unsafe output run inline on every turn, not as an afterthought. Tools and code execute in policy-governed sandboxes. And because the whole control plane is versioned, every change is shadow-tested and reversible — the opposite of a vendor deprecating a model under you. Capability is not the trade-off. Edge-native does not mean "smaller answers." Hard requests still reach a large on-prem model, and the gateway can escalate to a frontier API by explicit, attributed exception — so you keep the ceiling while removing the floor of wasted spend. 11 Recommendations & checklist Inventory steady-state workloads. Repetitive, high-volume, latency- or sovereignty-sensitive traffic is the first to move edge-native. Buy capacity, not tokens, for that traffic. Size a node fleet to steady demand; keep a frontier path for the long tail. Make the gateway the single ingress. One routing contract, one policy plane, one meter, compatible APIs so apps don't change. Turn on all four levers. Routing, prefix-cache discipline, context selection and semantic caching compound. Keep data on-prem by default. Treat egress as an exception that needs justification. Version and shadow-test policy. Never let routing drift silently; always be able to revert. Attribute cost continuously. Per-route metering turns AI spend into a capacity-planning input. Plan the two-tier fleet. Develop on the large node; serve on small nodes; promote artifacts unchanged. 12 References NVIDIA (2026). DGX Spark — Personal AI Supercomputer (GB10 Grace Blackwell). 128 GB LPDDR5X coherent unified memory; ~1,000 TFLOPS FP4; high-speed fabric for linked-pair scaling. AMD (2026). Ryzen AI Max+ 395 (Strix Halo). 16 Zen 5 cores, Radeon 8060S iGPU, XDNA 2 NPU; ~126 platform AI TFLOPS; 128 GB LPDDR5X-8000. Workload–router–pool architecture for inference optimization (2026). Signal-driven routing across a mixture of models by cost, capability, privacy and risk. When-to-reason routing (2025). Invoking expensive reasoning paths only when expected value justifies the cost. Category-aware semantic caching for heterogeneous workloads (2025). Reusing answers for semantically-equivalent requests. Prefix / KV-cache reuse in high-throughput serving. Cached prompt prefixes billed at a steep discount versus recomputation. Inference-native agent harness practice (2026). Prefix-cache discipline, context selection and bounded tool output for long-horizon work. Unovie.AI. GPU EdgeGateway & Device Platform. unovie.ai/platform/gpu-edgegateway · unovie.ai/device-platform Unovie.AI · EdgeAI Context Engineering. This whitepaper synthesizes internal edge-deployment practice with public hardware specifications and well-established inference-optimization concepts (signal-driven routing, prefix / KV-cache reuse, semantic caching, context selection). Cost figures are illustrative parameters, not price quotes; actual results depend on workload mix, utilization, token profile and energy cost. Hardware specifications reflect vendor figures as of June 2026. Revision 1.0.
# Modernizing the SOC for the Agentic Era — Edge-Native, Identity-Driven Security for Distributed IT/OT
URL: https://unovie.ai/ai-soc-modernization-whitepaper.html
## Modernizing the SOC for the Agentic Era: Edge-Native, Identity-Driven Security for Distributed IT/OT
Modernizing the SOC for the Agentic Era — Edge-Native, Identity-Driven Security for Distributed IT/OT ← Unovie.AI Download PDF Technical Whitepaper · AI SOC Modernization Modernizing the SOC for the Agentic Era: Edge-Native, Identity-Driven Security for Distributed IT/OT For CISOs and security organizations running distributed estates across IT and OT. Why the cloud-SIEM model breaks as agents — and agentic threats — appear everywhere, and how an edge-native, identity-driven AI SOC keeps detection, response and governance on your own ground. Suresh Mandava · suresh@unovie.com Unovie.AI · EdgeAI Context Engineering Revision 1.0 · June 2026 Reference hardware: NVIDIA Jetson AGX Thor · DGX Spark (Blackwell, 128 GB unified) — on-site Abstract The Security Operations Center evolved for two decades — NOC, SOC, AI/ML-assisted detection, threat intelligence, SOAR — toward a cloud-SIEM, managed-service model. That model assumed telemetry could leave the site and that threats were authored by humans. Both assumptions break in 2025 and beyond : AI agents now act everywhere — yours automating operations, adversaries' automating attacks — and distributed IT/OT estates cannot ship sensitive, high-volume telemetry to a metered cloud. We argue for an edge-native, identity-driven AI SOC : a self-learning reasoning core that runs on-site on owned silicon, detecting in the ingestion path; a knowledge graph that replaces isolated alerts with traced blast-radius; and an identity-driven core that gives every human, service and agent a scoped identity and policy — connected or air-gapped. Unovie delivers this through two offerings — GPU EdgeGateway , which governs agentic AI traffic, and the AI SOC (AISOC) , which detects and responds on your floor — bound by a portable identity edge. The result: detections that compound on your environment, response that survives outages, compliance evidence that is continuous, and data that never leaves your boundary. Contents The 2025+ inflection: agents everywhere Why the cloud-SIEM model breaks for IT/OT From 20 tools to four cognitive platforms The self-learning, edge-native SOC analyst From alerts to a knowledge graph The identity-driven core Governing agentic workloads The Unovie offering: EdgeGateway + AISOC Deployment for the distributed enterprise Governance, compliance & reversibility Recommendations for the CISO References 1 The 2025+ inflection: agents everywhere The SOC has been climbing the same ladder for twenty years. In 2025 the ladder ended, and the ground changed. Each rung added capability to a fundamentally reactive posture: the network operations center became a security operations center; rules gained machine-learning assists; threat intelligence enriched alerts; and SOAR automated the runbooks. Useful, incremental, human-paced. The 2025+ inflection is not another rung — it is a change in who acts. Agents are now on both sides of the wire. Adversaries automate reconnaissance, exploitation and lateral movement at machine speed; defenders deploy their own agents to triage, hunt and remediate. The volume, velocity and autonomy of action all step up at once. Figure 1 — Two decades of incremental, reactive evolution meet an inflection: action becomes autonomous and bidirectional. The SOC must defend at machine speed against — and with — agents. The shift The SOC's job is no longer to collect logs and correlate them later. It is to reason over relationships in the ingestion path, act autonomously within guardrails, and govern an estate where humans, services and agents all take actions that must be attributed. 2 Why the cloud-SIEM model breaks for IT/OT The 2022–2025 reference design was a cloud-SIEM, managed-service model: ship everything to a regional SIEM, correlate centrally, bill per query. For a distributed IT/OT enterprise that model now works against you. Assumption of the cloud-SIEM model Reality for distributed IT/OT in 2026+ Telemetry can leave the site OT, IoMT, video and process data are large, sensitive and often regulated — egress is a data-exfil and compliance risk Cost scales gracefully Cloud-SIEM fees recur on every event and query; agentic volume multiplies them without a ceiling The link is always up Plants, depots, substations and vehicles operate through outages; a cloud round-trip stalls incident response Generic detections are enough Vendor rules miss your environment; OT protocols and device behavior need locally-learned models Detection is a search problem At machine speed, delayed search and scheduled correlation are too late — detection must happen at ingest The OT estate compounds every one of these: long-lived devices, brittle protocols, safety constraints, no patch window, and air-gapped or intermittently-connected enclaves where a cloud dependency is simply unavailable. A modern SOC for this world has to run where the data is born . 3 From 20 tools to four cognitive platforms The traditional SOC is twenty fragmented tools, each a console and a silo. The modern architecture consolidates them into four cognitive platforms feeding a single reasoning core that takes autonomous action. Figure 2 — The 2026++ SOC. Four telemetry domains feed four cognitive platforms, which feed one self-learning reasoning core that drives four autonomous actions — entirely on owned silicon, with no cloud dependency. Each legacy control does not disappear; it is absorbed into a platform and made cognitive. The point is consolidation of data flows , not just consoles: Control What it was What the reasoning core makes it CSPM Finds cloud misconfig & drift Generates and tests infrastructure-as-code PRs to auto-close drift SIEM Centralized log search & retention Correlates events as a graph, not sequential log scans SOAR Hard-coded mitigation playbooks Playbooks reason over live topology instead of static scripts EDR Endpoint detect & remediate Explains endpoint anomalies and auto-scopes remediation MITRE ATT&CK Manual tactic mapping Auto-maps detections to adversary tactics in real time MEC (edge) Compute at the network edge Compiles hyper-local models that detect anomalies at the edge IDAM Identity, RBAC, MFA Adapts access policy to user, service and agent context in real time 4 The self-learning, edge-native SOC analyst The reasoning core is not a chatbot bolted onto a SIEM. It is a frozen open model adapted by external stores, improved by a verifier-graded loop, and run entirely on-device. Detection happens in the ingestion path , on the GPU, while events are still moving — tokenized, classified and enriched before indexing, using a compact transformer classifier rather than regex chains. A streaming bus in broker-only mode feeds parallel GPU workers; an inference server runs the model with dynamic batching; enriched incidents land in a sharded, authenticated index; a dead-letter queue protects failed batches and retention keeps storage bounded. On a single Blackwell-class node this sustains production-grade throughput: 21,300 + EPS Peak ingestion & AI-classification throughput, single node 13,800 + EPS Sustained production baseline ~3 s AI inference latency at ingest, under load 1.29 B/day Events at the 15K-EPS baseline (~1.55 TB enriched) Frozen base, reversible adaptation The model learns your environment without drifting. A frozen base never has its weights merged; adaptation lives in external, reversible stores — a knowledge layer (graph + retrieval), composable skills, and lightweight runtime controllers. A verifier-graded loop proposes updates, grades them against schema and grounding, and a regression gate commits only changes that beat the prior baseline on held-out data — otherwise it auto-reverts. Dual-path serving keeps a fast path for real-time detection and a deeper path for reflection. Why it is safe Because the base is frozen and every self-update is reversible and must clear an automatic regression gate before it goes live, detections only ever improve — the model compounds accuracy on your attacks, on-site, at near-zero marginal cost, with no weight drift and no data egress. 5 From alerts to a knowledge graph Isolated alerts hide multi-stage attacks. A knowledge graph models the enterprise as a web of relationships, so a low-severity signal can be traced to its true blast radius. Figure 3 — Relationship semantic context. When a low-severity anomaly fires on a workload, the core queries the graph — who owns the identity, what it talks to, where it runs, what it exposes — and traces the blast radius to catch the multi-stage attack the isolated alert would have missed. 6 The identity-driven core Distributed IT/OT and agentic workloads share one root requirement: every actor needs a verifiable identity and a scoped policy — whether the site is connected or air-gapped. That is the job of a thin identity edge. Rather than operate a heavy identity provider at every site, the architecture uses a thin authentication edge (an OAuth2/OIDC proxy with server-side sessions and edge RBAC) that delegates identity to the right issuer and injects a consistent identity context downstream. One pluggable setting selects the issuer; everything behind it stays identity-agnostic. Connected sites Air-gapped / disconnected enclaves Issuer The enterprise IdP — SSO, MFA and lifecycle stay where they already live; no local IdP to run A self-contained on-prem OIDC issuer with its own user/group store — no external database, no cloud reach Edge Same thin proxy; same server-side sessions; same RBAC on a roles claim; downstream services receive identity via standard headers Switch One pluggable issuer setting — applications and the SOC never change Secrets Held in the platform secret store; never committed; sessions in a local cache Identity-driven core A portable identity edge gives humans, services and agents one scoped identity model across the whole distributed estate — connected or air-gapped — and becomes the control point the SOC and the gateway both reason over. 7 Governing agentic workloads Agents are a new class of actor. They act on behalf of users, call tools, read sensitive data and spawn sub-agents — at machine speed. Each action must carry an identity, a scoped credential, a policy and an audit trail. This is where identity, the gateway and the SOC meet. The GPU EdgeGateway governs the agent's traffic : it routes each request to the right model on owned silicon, runs tools and code in policy-governed sandboxes (no unauthorized file, credential or network access), checks for sensitive-data leakage and prompt injection inline, and meters every call. The AI SOC governs the agent's behavior : it watches agent actions the way UEBA watches users, and uses the knowledge graph to bound what a compromised or misbehaving agent can reach. Agentic risk Control An agent acts with no attributable identity Identity-driven core issues a scoped identity per agent and human-on-behalf-of relationship An agent over-reaches its tools or data EdgeGateway sandboxes tools/code and enforces policy-as-code on every call A prompt injection or leak rides the request Inline safety classifiers (PII, jailbreak, injection) on every turn at the gateway A compromised agent moves laterally AISOC traces blast radius on the graph and contains by relationship Spend and action go unaudited Every route metered and attributed; every action logged for review 8 The Unovie offering: EdgeGateway + AISOC Unovie delivers this architecture as two offerings bound by the identity-driven core, all on hardware you own. GPU EdgeGateway — govern the agentic AI An inference-native, agent-first gateway. One routing contract turns signals into decisions across a mesh of local, private and frontier models; the prefix cache is protected; context is selected, not pasted; tools run sandboxed; and every policy change is shadow-tested before it goes live. It is the safe, governed, least-cost path for every agentic request — and the enforcement point for agent identity and policy. AI SOC (AISOC) — detect and respond on your floor An edge-native, self-learning security operations capability: GPU-native detection in the ingestion path, a knowledge graph for relationship context and blast-radius, a verifier-graded learning loop that compounds accuracy on your environment, and autonomous action — remediation-as-code, graph-traced containment, adaptive policy and continuous compliance evidence. It runs on-site, learns from your own attacks, and never ships telemetry off the boundary. See it on the platform. The detection-at-ingest, knowledge-graph and IT/OT coverage described here are productized as IT/OT Edge Security Intelligence — unovie.ai/platform/edge-security-intelligence . The agent-traffic governance layer is GPU EdgeGateway — unovie.ai/platform/gpu-edgegateway . 9 Deployment for the distributed enterprise An edge SOC node per site. Place a Blackwell-class node (Jetson AGX Thor / DGX Spark) where the data is — plant, depot, substation, clinic, vehicle bay — running detection-at-ingest and the reasoning core locally. One pipeline for IT and OT. IT telemetry (logs, identity, network, cloud) and OT/device telemetry land on the same local pipeline, so anomalies on the floor are correlated beside IT threats. Connected or air-gapped, same design. The identity edge points at the enterprise IdP where connected and a self-contained issuer where isolated; nothing downstream changes. Local autonomy, central correlation. Each node detects and responds on its own; a thin central layer correlates across sites without ever pulling raw telemetry off-site — only enriched, scoped findings travel. Govern agents through the gateway. All agentic AI traffic flows through EdgeGateway for routing, sandboxing, safety and metering. Own the model, not a subscription. Capacity is owned silicon; marginal cost approaches electricity; there is no per-query SIEM meter and no egress bill. 100 % On-site — telemetry never leaves the boundary $0 /query No recurring cloud-SIEM or per-event fees 24/7 Detection & response survive link and provider outages Weekly Detections compound from your own attacks 10 Governance, compliance & reversibility Owning the detection and inference path is the strongest governance posture available to a distributed enterprise. Sensitive IT/OT and regulated data never leaves the site, so data-exfil and cross-border transfer risk are removed at the source. Controls map continuously to frameworks — NIST guidelines, sector standards such as FHIR for health data — producing continuous compliance evidence rather than point-in-time audits. Because the reasoning core's base is frozen and every self-update is reversible behind a regression gate, the security model never drifts out from under you — the opposite of a cloud detection set that changes on a vendor's schedule. A live asset inventory (CMDB) tracks device firmware, end-of-life and protocol risk across the IT/OT estate, and the knowledge graph keeps identity, asset and vulnerability context joined. Capability without exposure. Hard problems still reach a large on-prem model, and the rare request that genuinely needs a frontier model can take that path by explicit, attributed exception through the gateway — so you keep the ceiling while keeping the data home. 11 Recommendations for the CISO Move detection to where the data is born. Stand up an edge SOC node per site; reserve the cloud for cross-site, scoped correlation only. Consolidate toward four cognitive platforms. Collapse the twenty-tool sprawl into cloud/workload, DevSecOps, cyber-ops and identity/asset platforms feeding one reasoning core. Detect at ingest, correlate on a graph. Treat detection as an in-path engineering problem; model the estate as relationships to catch multi-stage attacks. Make identity the core. Deploy a thin, pluggable identity edge so connected and air-gapped sites share one scoped model — for humans, services and agents. Govern every agent. Route agentic traffic through a gateway that sandboxes tools, enforces policy-as-code, checks safety inline and attributes every action. Own the model. Prefer a frozen-base, self-learning core you operate over a per-query subscription whose detections you do not control. Keep it reversible and provable. Gate every self-update on a regression test; map controls continuously to your compliance frameworks. Plan for IT and OT together. One pipeline, one identity model, one graph — not two disconnected programs. 12 References Unovie.AI. IT/OT Edge Security Intelligence. unovie.ai/platform/edge-security-intelligence — GPU-native detection at ingest, knowledge-graph correlation, IT + OT coverage. Unovie.AI. GPU EdgeGateway. unovie.ai/platform/gpu-edgegateway — inference-native, agent-first routing and governance on owned silicon. GPU-native SIEM reference architecture (single Blackwell node). Detection-at-ingest with a compact transformer classifier; ~21,300 EPS peak / ~13,800 EPS sustained; broker-only streaming, sharded authenticated index, dead-letter queue, retention enforcement. Self-learning, edge-native analyst pattern. Frozen base + external adaptation stores (knowledge / skills / controllers); verifier-graded loop with an automatic regression gate; dual-path serving; fully on-device. Thin authentication edge pattern. OAuth2/OIDC proxy with server-side sessions and edge RBAC; pluggable issuer — enterprise IdP (connected) or a self-contained on-prem OIDC issuer (air-gapped); identity-agnostic downstream. MITRE ATT&CK. Adversary tactics & techniques mapping used for real-time detection-to-tactic correlation. NIST guidance (incl. SP 800-190, container/IoT security) and sector standards (e.g., FHIR for health data interoperability). Continuous control mapping for compliance evidence. NVIDIA. Jetson AGX Thor & DGX Spark (Blackwell). On-site, 128 GB unified memory, NVFP4 — the owned-silicon target for edge SOC nodes. Unovie.AI · EdgeAI Context Engineering. This whitepaper synthesizes internal edge-security architecture with public hardware specifications and well-established security concepts (detection-at-ingest, knowledge-graph correlation, frozen-base self-learning, thin identity edges, MITRE ATT&CK, NIST). Throughput figures reflect a single-node reference benchmark and will vary with workload, hardware and configuration. Revision 1.0.
# Smart Factory Floor — Unovie.AI
URL: https://unovie.ai/solutions/smart-factory-floor.html
## Smart Factory Floor
Home / Solutions / Smart Factory Floor Manufacturing Smart Factory Floor Vision-driven inspection and quality control that runs on the line, in real time, on hardware you own — no cloud round-trip, no data egress. Start a project → Read the field guide −30 % inspection errors +25 % line speed 100 % on-prem
## Inspection & quality, automated
01 — What it does Inspection & quality, automated /vision Real-time defect detection Sub-second vision on the line catches defects humans miss at speed. camera→GPU sub-second /grading Consistent QC grading Eliminates inspector drift with a model that grades every unit identically. repeatable auditable /closed-loop Stop-the-line alerts Closed-loop alerting with full traceability when a defect threshold trips. alerting trace
## From frame to action
02 — How it works From frame to action 01 Capture Cameras stream to the on-prem GPU. 02 Detect Edge model flags defects per frame. 03 Decide Grade, route, or stop the line. 04 Log Every decision recorded for audit.
## Inspect every unit. Perfectly.
Let's build Inspect every unit. Perfectly. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Smart Warehouse — Unovie.AI
URL: https://unovie.ai/solutions/smart-warehouse.html
## Smart Warehouse
Home / Solutions / Smart Warehouse Logistics Smart Warehouse Edge-AI for inventory and movement: see stock in real time, optimize slotting, and cut stockouts — running where the cameras and sensors already are. Start a project → Read the field guide −20 % stockouts +15 % fill rate real-time tracking
## Inventory you can actually see
01 — What it does Inventory you can actually see /vision Inventory vision Continuous counts and location from existing cameras. counts location /slotting Slotting optimization Recommends placement to shorten pick paths. pick-path yield /trace Track & trace End-to-end movement with tamper-evident logs. trace audit
## Sense to optimize
02 — How it works Sense to optimize 01 Ingest Cameras & sensors → edge. 02 Model Count, locate, predict. 03 Optimize Slotting & replenishment. 04 Act Alerts & work orders.
## Never run out. Or over.
Let's build Never run out. Or over. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# OSHA Compliance Monitoring — Unovie.AI
URL: https://unovie.ai/solutions/osha-compliance.html
## OSHA Compliance Monitoring
Home / Solutions / OSHA Compliance Monitoring Safety OSHA Compliance Monitoring AI-powered safety analytics that spot compliance risks — missing PPE, unsafe zones, hazards — and alert in real time, with an evidence trail. Start a project → Read the field guide −25 % OSHA fines +30 % training done real-time alerts
## Safety, watched continuously
01 — What it does Safety, watched continuously /ppe PPE & hazard detection Detects missing protection and unsafe behaviour on camera. PPE zones /risk Risk analytics Trends incidents and near-misses into actionable risk. trends near-miss /evidence Alerting & evidence Real-time alerts plus a defensible audit log. alerting audit
## Monitor to audit
02 — How it works Monitor to audit 01 Monitor Watch zones & PPE. 02 Detect Flag risk events. 03 Alert Notify supervisors. 04 Record Immutable evidence log.
## Protect people. Prove it.
Let's build Protect people. Prove it. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Vision Intelligence — Unovie.AI
URL: https://unovie.ai/solutions/vision-intelligence.html
## Vision Intelligence
Home / Solutions / Vision Intelligence OTT · Content Provider Vision Intelligence Real-time computer vision for an OTT and pay-TV operator. Edge-AI watches every channel feed, identifies what is actually on screen frame by frame, and turns thousands of live streams into a continuously verified, audit-ready signal — on GPUs at the edge, in under a second. Start a project → Read the field guide 96 channels per node, live <1 s feed to verified −90 % manual QA effort
## See every stream, automatically
01 — What it does See every stream, automatically /detect Live channel detection GPU vision identifies the channel on every tile of every composite feed, with confidence and bounding box per frame. per-frame confidence bbox /verify Continuous QA & compliance Confirms each stream carries the right content and flags drift or outage in sub-second time, with annotated evidence for audit. MTTD↓ evidence audit /scale Many feeds, one node A zero-copy GPU pipeline decodes and scores dozens of 4K feeds per edge node — no raw video leaves the rack. zero-copy 4K on-prem
## Stream to verified signal
02 — How it works Stream to verified signal 01 Decode RTSP feeds decode into GPU memory. 02 Locate Per-tile regions mapped once. 03 Detect GPU scores each region every frame. 04 Publish Detections stream to analytics & alerts.
## Verify every channel, in real time.
Let's build Verify every channel, in real time. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Batch Process Optimization — Unovie.AI
URL: https://unovie.ai/solutions/batch-optimization.html
## Batch Process Optimization
Home / Solutions / Batch Process Optimization Process · chem & pharma Batch Process Optimization Automate the optimization of batch processes in chemical and pharmaceutical manufacturing — higher quality, less waste, per batch. Start a project → Read the field guide +25 % product quality −15 % waste per-batch tuning
## Tune every batch
01 — What it does Tune every batch /recipe Recipe optimization Recommends set-points from historical & live data. set-points yield /drift Drift detection Catches process drift before it costs a batch. drift SPC /yield Yield analytics Attributes yield to controllable factors. yield RCA
## Model to verify
02 — How it works Model to verify 01 Model Learn the process envelope. 02 Recommend Optimal set-points. 03 Apply Operator-in-the-loop. 04 Verify Score against held-out.
## More yield. Less waste.
Let's build More yield. Less waste. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Remote Expertise Platform — Unovie.AI
URL: https://unovie.ai/solutions/remote-expertise.html
## Remote Expertise Platform
Home / Solutions / Remote Expertise Platform Workforce Remote Expertise Platform Connect remote experts with on-site teams for real-time, grounded guidance — so knowledge scales without the travel. Start a project → Read the field guide +20 % worker output −30 % on-site visits real-time guidance
## Expertise, on demand
01 — What it does Expertise, on demand /guide Guided AR / voice Step-by-step guidance overlaid where the work happens. AR voice /ground Grounded retrieval Answers cited to your manuals and records. RAG cited /capture Session capture Turns each session into reusable knowledge. capture reuse
## Connect to capture
02 — How it works Connect to capture 01 Connect Link expert & floor. 02 Ground Pull cited context. 03 Guide Real-time steps. 04 Capture Bank the knowledge.
## Your best expert, everywhere.
Let's build Your best expert, everywhere. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Maritime Digital Twin — Unovie.AI
URL: https://unovie.ai/solutions/maritime-digital-twin.html
## Maritime Digital Twin
Home / Solutions / Maritime Digital Twin Maritime · Oil & Gas Maritime Digital Twin A living digital twin for large tanker and LNG fleets. Edge-AI on every vessel watches vibration, acoustics, hull strain and the OT network, classifies anomalies before the satellite uplink, and turns condition evidence into deferred-maintenance and compliance decisions on shore. Start a project → Read the field guide −6 mo overhauls deferred on evidence 100 % classified on-vessel 800 vessels · one twin
## Intelligence at the waterline
01 — What it does Intelligence at the waterline /edge Edge anomaly detection Vibration, acoustic and strain models run on the vessel and classify faults before the uplink — so low-bandwidth links carry decisions, not raw signal. vibration acoustic strain /cbm Condition-based maintenance Remaining-useful-life on rotating gear becomes defensible deferral evidence for class surveys, and auto-triggers spares procurement. RUL deferral spares /cyber OT cyber watch An on-network sensor flags anomalous engine and automation commands, isolates the endpoint, and keeps a full forensic trail. OT IDS isolate forensics
## From sensor to shore decision
02 — How it works From sensor to shore decision 01 Sense High-rate sensors on engine, hull and OT bus. 02 Classify Edge models score events on the vessel. 03 Sync Only decisions cross the satellite link. 04 Act Defer, procure, or isolate — with evidence.
## Run the fleet on evidence.
Let's build Run the fleet on evidence. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Connected-Vehicle Digital Twin — Unovie.AI
URL: https://unovie.ai/solutions/connected-vehicle-twin.html
## Connected-Vehicle Digital Twin
Home / Solutions / Connected-Vehicle Digital Twin Automotive · Connected EV Connected-Vehicle Digital Twin A VIN-level digital twin for multi-brand connected-EV fleets. As-ordered, as-built, as-operated and as-maintained data unify into one live view, with an in-vehicle AI assistant and battery diagnostics running on automotive-grade edge silicon and delivered over the air. Start a project → Read the field guide 4.2 h average case resolution 98.7 % notification delivery 4 lifecycle states · one VIN
## One VIN, every dimension
01 — What it does One VIN, every dimension /twin Lifecycle digital twin As-ordered, as-built, as-operated and as-maintained unified per VIN — live telemetry beside service, recall and warranty history. telemetry BOM history /assist In-vehicle AI assistant A generative assistant and battery state-of-health diagnostics run on automotive edge SoCs, with new models pushed over the air. GenAI battery SoH OTA /care Context-aware customer care Inbound contacts pop full vehicle context, next-best-action and consent state; video remote-assist and digital-key control close the loop. screen-pop NBA digital keys
## Signal to service
02 — How it works Signal to service 01 Connect Vehicles stream telemetry and faults. 02 Diagnose Edge models score battery and systems. 03 Enrich Cases pop context and next action. 04 Resolve Remote command, dispatch, or OTA.
## Know every vehicle, by VIN.
Let's build Know every vehicle, by VIN. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Corporate-Travel Sales Intelligence — Unovie.AI
URL: https://unovie.ai/solutions/corporate-travel-sales.html
## Corporate-Travel Sales Intelligence
Home / Solutions / Corporate-Travel Sales Intelligence Go-to-market · Airline Corporate-Travel Sales Intelligence An AI co-pilot for an airline corporate-sales team. It grounds each corporate account in live travel-demand signals, scores fit across seven airline-specific dimensions — network overlap, premium-cabin propensity, loyalty, travel-policy maturity and more — drafts a full opportunity plan with stakeholder map and competitive positioning, and surfaces the next best action across the whole book of business. Start a project → Read the field guide <5 min meeting prep, from an hour 1 -day RFP turnaround, from five 2,000 + accounts, one book
## From signal to corporate deal
01 — What it does From signal to corporate deal /signal Travel-demand signal grounding Hiring, expansion, earnings and RFP signals per account distil into issue, impact and opportunity a rep can act on. grounded signals RAG /score Airline-specific fit scoring Seven weighted dimensions — network overlap, premium-cabin propensity, loyalty, policy maturity, buying intent and more — render an instant fit and propensity picture. 7 axes propensity explainable /leak Share-of-wallet leakage Route-level share against contracted commitment flags slipping accounts and the revenue at risk — before renewal. share routes at-risk
## Account to next move
02 — How it works Account to next move 01 Ground Pull live demand signals. 02 Score Rank fit across seven axes. 03 Plan Draft the opportunity and stakeholders. 04 Act Surface the next best action.
## Win the corporate account.
Let's build Win the corporate account. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Edge Data Fabric — Unovie.AI
URL: https://unovie.ai/platform/edge-data-fabric.html
## Edge Data Fabric
Home / Platform / Edge Data Fabric Platform · Nexus Edge Data Fabric Unovie helps you own your enterprise context foundation — on your terms, with your data — to power the AI-native initiatives ahead. A domain ontology, a knowledge graph and vector memory become one living model of your world: you can buy the models and the agents, but the context — the boundary that makes you a coherent system — is yours to build, woven from open standards, on-prem. Custom models read your documents, telemetry and records into a typed graph; agents reason over it; your teams query it in plain language. And leaked context — vectors handed to models you do not own — is a business liability and a governance violation. Start a project → Read the field guide ontology domain-modelled graph +vector hybrid recall 0 B egress · on-prem
## Context that compounds
01 — What it does Context that compounds /ontology A domain ontology Your world modelled as typed entities and relationships — the who, what, when, where and why — so context is structured, not guessed. typed relationships semantic /graph Knowledge graph + vector A graph store and vector memory side by side: subgraph traversal for structure, dense and sparse embeddings for meaning, fused into one answer. graph vector hybrid /selfserve Self-service by language Teams author and query in natural language; domain fine-tuned agents plan, retrieve and act — no SQL, no data team in the loop. NL query agents no-code
## Documents in, decisions out
02 — How it works Documents in, decisions out 01 Extract Parse docs, tables and signals. 02 Construct Build the typed graph. 03 Index Embed for graph + vector. 04 Serve Grounded answers & agents.
## Your context is your membrane.
The context membrane Your context is your membrane. Your context is not a feature you bolt on — it is your model of your own world, the boundary that lets your organisation perceive, predict and act as one coherent system. An ontology core, a living knowledge graph, and a membrane of identifiers, rules, meaning and processes that lets the world in without losing what makes you you. A context you can buy is a context your competitors can buy too. Leaked context — or vectors — to models you do not own is a business liability. A governance violation. You can buy the models. You can buy the agents. The context — the boundary that defines you — you build and own, woven from open standards, on your terms. Outsource it, and you become a component in a system someone else defines.
## How the graph is built
03 — Architecture How the graph is built /parse Custom extraction models Document- and table-structure models parse PDFs, images and text; an LLM emits structured records that become typed nodes and edges. doc-parse tables LLM /construct Graph construction + embeddings Records are de-duplicated, transformed and persisted to a fast embeddable graph store, with embedding models writing vectors beside every node. KG build embeddings transform /retrieve Hybrid fused retrieval Traditional indices, text and sparse embeddings and subgraph traversal feed a tensor-based fused ranker and query-rewrite models — grounded Top-K, not guesses. fusion rerank query-rewrite /version Temporal & governed Graph OLTP keeps every node and edge temporally versioned with an immutable audit trail; columnar OLAP serves analytics — multi-tenant, OAuth2 / RBAC, zero-trust. temporal audit RBAC
## Models and agents, tuned to your domain
04 — Custom models Models and agents, tuned to your domain /models A catalog of custom models Document-parsing, embedding, NER and query-rewrite models, plus domain fine-tuned SLMs — trained on your domain, swappable and multi-model routed. fine-tuned SLM NER multi-model /agents Domain agent accelerators Pre-built agents — diagnostics, differential, referral, next-best-action and triage — reason over the graph and call tools to act. agents NBA tools /nlp Natural-language self-service Self-service portals let users ask, author and govern in plain language across chat, voice and app — every answer traceable to source. NL authoring omni-channel traceable
## Built for scale
05 — By the numbers Built for scale 100 M object nodes 2.5 B relationships 260 + prebuilt connectors
## Know what you know — and why.
Let's build Know what you know — and why. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Edge Streaming Intelligence — Unovie.AI
URL: https://unovie.ai/platform/edge-streaming-analytics.html
## Edge Streaming Intelligence
Home / Platform / Edge Streaming Intelligence Platform Edge Streaming Intelligence Real-time vision and audio intelligence over live streams, at fleet scale. A zero-copy GPU pipeline decodes hundreds of feeds straight into device memory, runs a catalog of detectors on every frame, and turns raw video into a verified, queryable signal — with sub-millisecond validation and an agent layer that reasons and acts on top. Proven on broadcast-grade media: 1,000+ concurrent 4K streams across racks of edge GPUs. Start a project → Read the field guide 1,000 + 4K streams, concurrent <1 ms high-frequency validation 128 streams / rack
## Vision and audio on every frame
01 — What it does Vision and audio on every frame /detect A catalog of live detectors Logos, freezes, macro-blocking, blank and splash screens, lip-sync and on-screen errors are scored per frame — video and audio anomalies caught the instant they appear. video audio per-frame /read Reads the screen, not just watches it OCR and vision-language models extract guide data, clocks, version strings and error dialogs; object detectors track focus, icons and UI state. OCR/VLM object-det UI state /act Closes the loop An agent layer plans and acts — driving devices through an IR / Bluetooth control plane and verifying every step against the live stream. agentic device-control verify
## Stream to decision
02 — How it works Stream to decision 01 Decode Feeds decode into GPU memory. 02 Detect A model graph scores every frame. 03 Decide Validate sub-ms; reason in minutes. 04 Act Drive devices, publish, alert.
## From satellite to every screen.
Live streams From satellite to every screen. Telecom and satellite feeds land at the edge, get scored frame by frame, and serve verified media — at fleet scale.
## A model graph, not a single model
03 — Detection catalog A model graph, not a single model /anomaly Video anomaly detection Freezes (consecutive pixel-difference), macro-blocking and pixelation (block-variance + Sobel edge density), tearing and stutter — flagged inside the stream buffer. optical-flow Sobel block-variance /logo Logo & UI object detection An RF-DETR detector with a CLIP refiner confirms logos, app tiles and widgets with bounding-box precision. RF-DETR CLIP bbox /ocr OCR & VLM reading GPU OCR (docTR) and vision-language models read guide grids, clocks, version strings and error dialogs — signal-loss, auth and tune failures included. docTR VLM regex /audio Audio & sync checks Audio-presence and lip-sync checks run beside the video probe, so silent feeds and A/V drift are caught too. audio-probe lip-sync ffprobe
## Inside the pipeline
04 — Architecture Inside the pipeline /pipeline Zero-copy vision pipeline GStreamer + DeepStream pull RTSP / H.265 into the GPU via NVDEC; composite grids map to regions once, then a swappable model graph scores each region every frame. GStreamer DeepStream NVDEC /serverless Serverless model serving Detectors run as auto-scaling GPU functions (Nuclio) drawn from a continuously trained catalog — new models deploy without touching the pipeline. Nuclio auto-scale registry /backbone Event & knowledge backbone Detections stream over NATS JetStream into ClickHouse for sub-second OLAP, with a knowledge graph, vectors and a fine-tuned vision-action model driving next-best-action. NATS ClickHouse knowledge-graph /learn A closed training loop Misses become flagged frames become new annotation tasks — captured, versioned in COCO and retrained, then promoted through a registry. CVAT COCO feedback
## Engineered for density
05 — By the numbers Engineered for density ~512 streams · 4 racks in parallel <50 ms inference / frame 30 fps sustained per stream
## Real-time, actually.
Let's build Real-time, actually. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# GPU MicroCloud — Unovie.AI
URL: https://unovie.ai/platform/gpu-microcloud.html
## GPU MicroCloud
Home / Platform / GPU MicroCloud Platform GPU MicroCloud We stand up a full private cloud on your floor: racks of edge GPUs and servers, software-defined storage, a 10G fabric and a Kubernetes control plane — all on-prem. Workloads are scheduled and bin-packed across the pool, MIG carves each GPU into isolated slices, zero-trust policy gates every call, and every minute is metered for chargeback. Datacenter discipline, where your data lives. Start a project → Read the field guide Kubernetes cloud-native, on-prem MIG hard isolation metered per-tenant chargeback
## Your GPUs, run like cloud
01 — What it does Your GPUs, run like cloud /schedule Fair-share scheduling A priority-and-quota scheduler places jobs across the pool, preempts politely, and keeps expensive silicon busy. queue quota preempt /isolate Hard multi-tenant isolation MIG partitioning carves each GPU into isolated slices, so tenants share hardware without sharing blast radius. MIG cgroups secure /meter Metering & chargeback Per-tenant, per-job accounting turns shared capacity into auditable cost and showback reports. metering showback reports
## Pool to bill
02 — How it works Pool to bill 01 Pool Aggregate edge GPUs. 02 Schedule Place workloads. 03 Isolate Partition tenants. 04 Meter Account & bill.
## Many nodes, one intelligence.
Distributed by design Many nodes, one intelligence. GPU nodes pool as one mesh — scheduled, partitioned and sharing state through a control plane. Workloads land anywhere; the cloud behaves as one.
## Inside the micro-cloud
03 — Architecture Inside the micro-cloud /compute Heterogeneous compute pool Edge GPUs, CPU servers and training boxes are pooled as one schedulable fabric — production racks plus a dedicated test / stage rack. edge GPU servers DGX /orchestrate Kubernetes control plane A cloud-native control plane handles dynamic model deployment, job scheduling, capacity allocation and execution failover across namespaced dev / test / prod. Kubernetes Helm failover /storage Software-defined storage CEPH block / file / object, an S3-compatible object store and NFS / iSCSI SAN give every workload durable, shared state — no external cloud. CEPH S3 NFS/iSCSI /partition MIG slice fabric Each GPU is partitioned into right-sized MIG instances and bin-packed by memory and NVLink topology — tenants share silicon, never blast radius. MIG NVLink bin-pack
## Governed like a cloud region
04 — Governance Governed like a cloud region /zerotrust Zero-trust by default Policy-as-code and OAuth2 / OpenID gate every call; per-tenant namespaces and RBAC separate workloads and data end to end. OPA OAuth2/OIDC RBAC /observe Deep observability OpenTelemetry and eBPF trace every workload; metrics, logs and dashboards make utilization and cost visible in real time. OpenTelemetry eBPF Grafana /data Stateful data services Cache, queue and database services run inside the cloud beside the compute, with columnar analytics for usage and reporting. Redis PostgreSQL ClickHouse /meter Metering & chargeback Per-tenant, per-job accounting turns shared capacity into auditable showback — quotas, reports and capacity planning. metering quotas showback
## The fabric underneath
05 — By the numbers The fabric underneath 10 G overlay SAN + LAN 7 × MIG slices / GPU 3 namespaces · dev/test/prod
## Datacenter discipline, on-prem.
Let's build Datacenter discipline, on-prem. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# GPU EdgeGateway — Unovie.AI
URL: https://unovie.ai/platform/gpu-edgegateway.html
## GPU EdgeGateway
Home / Platform / GPU EdgeGateway Platform GPU EdgeGateway Most gateways sit in front of the models and treat inference as a black box. This one is built from the inference engine out: one routing contract turns signals into projections, projections into decisions, and decisions into the model — across a mesh of local, private and frontier engines — while the prefix cache is protected, context is selected rather than pasted, and every turn takes the least-cost path that still meets the need. Session-aware across long-running agents, sandboxed for tool safety, and shadow-tested before any policy goes live. OpenAI- and Anthropic-compatible, multimodal, governed like production, on hardware you own. Start a project → Read the field guide signal→model one routing contract prefix-cache reuse-protected least cost per-turn path
## Route, reason, act
01 — What it does Route, reason, act /route Signal-driven routing Intent, complexity, modality and risk become projections and policy bands, then route across a local-to-frontier mesh — reasoning only when it pays. intent projections when-to-reason /session Session-aware agentic routing Stateful guards keep multi-turn agents coherent: hard locks block unsafe model switches mid-tool-loop, weighing quality gap, prefix locality and turn priors. session-aware tool-loop locks continuity /sandbox Sandboxed & governed Tools and code run in policy-governed MicroVM sandboxes — no unauthorized file, credential or network access. MicroVM policy-as-code no-exfil
## Signal to decision
02 — How it works Signal to decision 01 Signal Score intent, risk, modality, context. 02 Project Normalise into policy bands. 03 Decide Pick model, agent or tool. 04 Serve Sandboxed, observed, metered.
## A self-improving router.
Control plane · data plane A self-improving router. A control plane governs policy, identity and guardrails; a data plane serves fast, observable, cost-aware inference; and a self-improving router between them turns every request into a better next decision — protecting the prefix cache and selecting context so the work stays cheap as it grows.
## Inside the gateway
03 — Architecture Inside the gateway /contract One routing contract Signals become projections, projections drive decisions, decisions choose the model — the same pipeline whether configured in YAML, the console, the CLI or Kubernetes. signals projections decisions /mesh Mixture-of-models mesh Token- and capability-aware routing spans self-hosted engines, local SLMs and frontier APIs with semantic caching; classifiers run on any accelerator — one control plane, any backend. self-hosted semantic cache any accelerator /safe Safety & protocol History-aware PII, jailbreak and prompt-injection scanning across every turn — behind an OpenAI- and Anthropic-compatible ingress with explicit, lossless translation. PII jailbreak OpenAI/Anthropic /cache Prefix-cache discipline Stable prompt epochs, deterministic tool-schema ordering and bounded, append-only context keep reusable prefixes intact — so cached tokens are reused across a long session at a fraction of the price instead of re-billed every turn. prompt epochs stable schema cache reuse /lifecycle Shadow, activate, revert Every routing policy is versioned and shadow-tested on replayed traffic before activation, with one-click rollback — routing never drifts silently. shadow replay rollback
## Multimodal in, action out
04 — Agent-first delivery Multimodal in, action out /multimodal Every modality, one path Text, voice, image and event inputs are normalised, routed to the right modality model, and turned into grounded responses or tool actions. text·voice·image normalise actions /context Context selected, not pasted Graph-shaped code evidence, bounded tool output and domain-aware compression extract the signal a turn actually needs and drop the rest — fewer prompt and tool-output tokens, without losing continuity across a long task. select not paste graph context bounded output /observe Topology & token ledger A console traces every signal → projection → decision with replay, and a live ledger shows cache reuse, context savings and per-route latency, tokens and cost — spend is accountable while the task runs, not after. topology savings ledger metering
## Governed like production
05 — By the numbers Governed like production <1 ms signal → decision ~90 % cached-token discount least cost per-turn path shadow→activate policy lifecycle
## Whitepapers for your team
06 — Further reading Whitepapers for your team The architecture and the economics behind this platform — read in the browser or export to PDF. /economics The Edge-Native Inference Gateway For infrastructure leaders: turning unpredictable, metered AI opex into fixed, predictable cost for the industrial edge. predictable cost edge-native PDF /doctrine Training Without Retraining The frozen-base doctrine: adapting custom models on the edge through context and self-verification, not weight updates. frozen-base on-device PDF
## Serve models safely.
Let's build Serve models safely. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# IT/OT Edge Security Intelligence — Unovie.AI
URL: https://unovie.ai/platform/edge-security-intelligence.html
## IT/OT Edge Security Intelligence
Home / Platform / IT/OT Edge Security Intelligence Platform IT/OT Edge Security Intelligence A GPU-native SIEM that detects threats while the data is still moving. Instead of collecting logs and correlating them later, it tokenizes, classifies and enriches every event in flight on the GPU — semantic AI detection, not regex chains — then indexes to a sharded, authenticated store. Tens of thousands of events per second on a single edge node, on-prem. Start a project → Read the field guide 21K + events / sec, peak semantic AI, not regex single-node on-prem
## Detection at ingest
01 — What it does Detection at ingest /inflight Detection in the data path Events are tokenized, classified and scored on the GPU as they stream in — alerts fire near ingestion, not after a delayed search job. in-flight low-latency streaming /semantic Semantic threat detection A BERT classifier reads intent and meaning in raw log text, catching threats that static rules and regex miss. BERT intent beyond-regex /resilient Hardened & bounded A dead-letter queue protects failed batches, retention keeps storage bounded, and authenticated, sharded indexing keeps search fast. DLQ retention auth
## Log to incident
02 — How it works Log to incident 01 Ingest Logs land in a Kafka stream. 02 Batch Workers tokenize on the GPU. 03 Classify BERT inference scores intent. 04 Index Enriched incidents to search.
## Threats stop at the edge.
Edge defense Threats stop at the edge. Five nested layers — business contexts at the core, then business data, edge agents and sensors — wrapped in one shield. DDoS floods and intrusions are detected and deflected before they ever reach the core.
## Inside the pipeline
03 — Architecture Inside the pipeline /ingest Streaming ingest High-throughput Kafka in KRaft mode (no ZooKeeper) feeds parallel consumers — backpressure-safe at tens of thousands of events per second. Kafka KRaft parallel /infer GPU inference server An inference server runs the detection model on the GPU in batches, so classification scales with parallelism instead of CPU cores. Triton Morpheus GPU batch /enrich Enrich & index Scores and metadata are attached, then incidents are written to an authenticated, multi-shard search index for fast investigation. enrichment Elasticsearch 8-shard /harden Operational hardening Dead-letter queue, health checks, retention enforcement and authentication keep the pipeline resilient and storage bounded. DLQ healthchecks retention
## One lens over both estates
04 — IT + OT coverage One lens over both estates /it IT telemetry Logs, endpoints, identity and network events are classified for intent — credential abuse, lateral movement and exfiltration patterns surfaced in flight. logs identity network /ot OT & edge signals Operational-technology and device telemetry are watched on the same pipeline, so anomalies on the plant floor and at the edge are caught beside IT threats. OT ICS device /correlate Unified incidents IT and OT detections land in one store with shared scoring and timelines — correlation across both estates, not two disconnected tools. correlation timeline single-pane
## Engineered for throughput
05 — By the numbers Engineered for throughput 21K + EPS peak 13.8K + EPS sustained ~3 s AI inference latency
## Whitepapers for your team
06 — Further reading Whitepapers for your team The architecture and the economics behind this platform — read in the browser or export to PDF. /ai-soc AI SOC Modernization For CISOs: an edge-native, identity-driven AI SOC for distributed IT/OT in the agentic era — detection at ingest, knowledge-graph context, autonomous response. IT/OT identity-driven PDF /economics The Edge-Native Inference Gateway Turning unpredictable, metered AI opex into fixed, predictable cost for the industrial edge. predictable cost edge-native PDF
## Detect threats in the data path.
Let's build Detect threats in the data path. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Device Platform — NVIDIA, Qualcomm and AMD edge AI — Unovie.AI
URL: https://unovie.ai/device-platform.html
## Hardware you own.
Home / Device Platform Device Platform Hardware you own. The NVIDIA, Qualcomm and AMD platforms we build, optimize and operate end-to-end — from a power-efficient far-edge processor to a desktop AI supercomputer. The same Unovie stack, your data, on silicon you own. Start a project → Read the field guide
NVIDIA AGX Thor NVIDIA DGX Spark Qualcomm QCS6490 AMD Ryzen AI Max+ 395
## NVIDIA AGX Thor
Device · Robotics edge NVIDIA AGX Thor A Blackwell-class edge supercomputer for physical AI. NVIDIA Jetson AGX Thor packs up to 2,070 FP4 TFLOPS of generative-AI compute and 128 GB of unified memory into a power-configurable module small enough to live inside a robot, a vehicle or a machine — running several large models, vision and multi-sensor fusion at once, on-prem. We build, optimize and operate the full Unovie stack on Thor, so your edge agents run where the data is born. 2,070 TFLOPS FP4 AI compute 128 GB unified LPDDR5X ~7.5 × vs Jetson Orin
## An edge box you can hold.
The device An edge box you can hold. Thor ships as a compact, fan-cooled edge node: a dense I/O wall of USB, networking, display and capture, with a Blackwell GPU and 128 GB of unified memory behind it. Mount it on the line, in the cab or at the cell — and run the models where the data is born.
## Physical AI at the edge
01 — What it does Physical AI at the edge /blackwell Blackwell on a module A datacenter-class Blackwell GPU with FP4 and a transformer engine, packed into a module — generative and vision models that used to need a rack now run inside the machine. Blackwell FP4 transformer-engine /fusion Multi-sensor, multi-model A 14-core Arm Neoverse CPU and high-bandwidth memory run camera, lidar, radar and language models together, fused in real time for autonomy and inspection. sensor-fusion multi-model real-time /safety Partitioned & safety-ready MIG carves the GPU into isolated slices inside a configurable 40–130W envelope, with a functional-safety design for robots and autonomous machines. MIG 40–130W safety
## Silicon to autonomy
02 — How it works Silicon to autonomy 01 Provision Image Thor with the Unovie edge stack. 02 Serve Local models + Nexus context, on-device. 03 Fuse Vision, sensors and agents reason live. 04 Act Closed-loop control, fully on-prem.
## Built for the machine
03 — Architecture Built for the machine /compute Blackwell GPU + Tensor Cores 2,560 CUDA cores and next-gen Tensor Cores with FP4 and a transformer engine for on-device generative AI. CUDA Tensor FP4 /cpu 14-core Arm Neoverse A 14-core Arm Neoverse-V3AE cluster feeds the GPU and runs the control plane, sensors and OS. Neoverse-V3AE 14-core /io Sensor-grade I/O High-speed camera, networking and PCIe lanes ingest many sensors at once with deterministic latency. MIPI/CSI PCIe 10/25G
## By the numbers
04 — By the numbers By the numbers 2,070 TFLOPS FP4 (sparse) 128 GB LPDDR5X 273 GB/s memory bandwidth 40–130 W configurable
## NVIDIA DGX Spark
Device · Desktop supercomputer NVIDIA DGX Spark A petaFLOP AI supercomputer that fits on a desk. NVIDIA DGX Spark pairs the GB10 Grace Blackwell Superchip with 128 GB of coherent unified memory and up to 1,000 TFLOPS of FP4 compute — enough to prototype, fine-tune and run models up to ~200B parameters locally, or ~405B across a linked pair. We run it as your private development and inference node: the full Unovie stack, your data, your room. 1,000 TFLOPS FP4 AI compute 128 GB coherent memory 200 B params, local
## A supercomputer that fits on a desk.
The device A supercomputer that fits on a desk. Spark is a desktop-sized chassis with a perforated cooling top and a full I/O wall — a GB10 Grace Blackwell Superchip and 128 GB of coherent memory inside. Develop, fine-tune and serve large models locally; promote them to the edge unchanged.
## A supercomputer you own
01 — What it does A supercomputer you own /gb10 Grace Blackwell GB10 A 20-core Arm Grace CPU and a Blackwell GPU joined by NVLink-C2C share one coherent memory space — no PCIe copies between CPU and GPU. GB10 NVLink-C2C coherent /memory 128 GB for big models Unified LPDDR5X holds models up to ~200B parameters; two units linked over ConnectX scale to ~405B — inference and fine-tuning without the cloud. 200B local 405B linked ConnectX /stack The full NVIDIA AI stack Runs NIM microservices, CUDA frameworks and the same containers as DGX in the datacenter — develop locally, deploy to the edge unchanged. NIM CUDA portable
## Desk to deployment
02 — How it works Desk to deployment 01 Build Prototype & fine-tune locally on Spark. 02 Ground Wire in your Nexus context and data. 03 Validate Run the same containers as production. 04 Promote Ship unchanged to edge or MicroCloud.
## One coherent memory space
03 — Architecture One coherent memory space /superchip GB10 Grace Blackwell Grace CPU and Blackwell GPU on one package, joined by NVLink-C2C at chip-to-chip bandwidth. GB10 NVLink-C2C /memory 128 GB unified LPDDR5X CPU and GPU address one coherent pool — no host-device copies, and room for ~200B-parameter models. unified coherent 200B /fabric ConnectX scale-out ConnectX networking links two Sparks into a single ~405B-parameter inference target. ConnectX RDMA 405B
## By the numbers
04 — By the numbers By the numbers 1,000 TFLOPS FP4 AI 128 GB unified memory 20 Arm Grace cores 4 TB NVMe storage
## Qualcomm QCS6490
Device · Power-efficient edge Qualcomm QCS6490 A power-efficient edge-AI processor for robots, cameras and handhelds. The Qualcomm QCS6490 pairs an octa-core Kryo CPU, an Adreno GPU and a Hexagon AI processor for up to 12 TFLOPS — multi-camera vision and on-device models on a fanless, battery-friendly power budget, with Wi-Fi 6E and long industrial lifecycle support. We bring the Unovie stack to it, so intelligence runs at the far edge, on hardware you own. 12 TFLOPS Hexagon AI 5 concurrent cameras Wi-Fi 6E FastConnect
## Built for the far edge.
The device Built for the far edge. QCS6490 reference hardware brings a full I/O wall — USB-C, USB 3.0, dual Ethernet, 10GbE and HDMI — to a compact, fanless box. Premium-tier on-device AI without the power bill, deployed where wires and watts are scarce.
## AI on a power budget
01 — What it does AI on a power budget /hexagon Hexagon AI at low watts Up to 12 TFLOPS from the Hexagon processor with a fused tensor accelerator — vision, speech and sensor models on a budget that fits a fanless box or a battery. 12 TFLOPS Hexagon low-power /vision Triple ISP, many cameras A Spectra triple ISP ingests up to five concurrent cameras with computer-vision hardware — multi-camera perception for robots, handhelds and smart cameras. Spectra ISP 5 cameras CV /connect Wi-Fi 6E, built to last FastConnect Wi-Fi 6E and Bluetooth 5.2 keep the edge connected wirelessly, with wide-temperature, long-lifecycle industrial availability. Wi-Fi 6E BT 5.2 industrial
## Sense to inference
02 — How it works Sense to inference 01 Capture Up to 5 cameras and sensors stream in. 02 Process Kryo CPU + Adreno GPU + Hexagon NPU. 03 Infer Vision and language models on-device. 04 Connect Results over Wi-Fi 6E, no cloud.
## A heterogeneous compute engine
03 — Architecture A heterogeneous compute engine /cpu Octa-core Kryo CPU A 6 nm octa-core Qualcomm Kryo CPU runs the OS, control and classical workloads beside the AI engines. Kryo octa-core 6 nm /npu Hexagon + Adreno The Hexagon processor with a fused tensor accelerator and the Adreno GPU share inference and graphics — up to 12 TFLOPS. Hexagon Adreno 12 TFLOPS /isp Spectra triple ISP A triple ISP captures up to five concurrent camera streams with 4K HDR video and on-sensor computer vision. Spectra 5 cameras 4K HDR
## By the numbers
04 — By the numbers By the numbers 12 TFLOPS Hexagon AI 5 concurrent cameras Wi-Fi 6E FastConnect 6 nm process
## AMD Ryzen AI Max+ 395
Device · Private AI server AMD Ryzen AI Max+ 395 A private AI server in a small metal box. The AMD Ryzen AI Max+ 395 fuses 16 Zen 5 CPU cores, a Radeon 8060S iGPU and a next-gen XDNA 2 NPU for 126 platform AI TFLOPS, paired with 128 GB of LPDDR5X-8000 — enough to run 70B-class models locally, behind dual 10GbE and USB4 so nodes cluster into a compute hub. We deploy the Unovie stack on it for secure, private inference on hardware you own. 126 TFLOPS platform AI 128 GB LPDDR5X-8000 70 B models, local
## A server that hides in plain sight.
The device A server that hides in plain sight. An all-metal chassis with a built-in 230 W supply exposes dual 10GbE, dual USB4 and fast PCIe 4.0 NVMe on its I/O wall — a quiet, durable node you can rack a few of, or set one on a desk.
## A private model server
01 — What it does A private model server /apu 16 Zen 5 + Radeon + XDNA 2 Sixteen Zen 5 CPU cores, a Radeon 8060S iGPU and a next-gen XDNA 2 NPU combine for 126 AI TFLOPS — CPU, GPU and NPU inference in one package. Zen 5 Radeon 8060S XDNA 2 /memory 128 GB for big models 128 GB of LPDDR5X-8000 keeps large models — 70B-class and up — resident and private, with no weights leaving the box. 128 GB LPDDR5X-8000 70B local /cluster Clusters into a hub Dual 10GbE and dual USB4 at 40 Gbps link nodes into an AI compute hub for distributed, local inference. dual 10GbE USB4 40G clustering
## Box to private cloud
02 — How it works Box to private cloud 01 Load 70B-class models resident in 128 GB. 02 Serve CPU + Radeon iGPU + XDNA 2 NPU. 03 Cluster Link nodes over 10GbE / USB4. 04 Operate Private inference, fully on-prem.
## One package, three engines
03 — Architecture One package, three engines /cpu 16 Zen 5 cores A 16-core Zen 5 CPU drives orchestration, data prep and classical workloads alongside inference. Zen 5 16-core /gpu Radeon 8060S + XDNA 2 The Radeon 8060S iGPU and XDNA 2 NPU share AI work for 126 TFLOPS across vision, language and agents. Radeon 8060S XDNA 2 126 TFLOPS /thermal 140W, vapor-chamber cooled Dual turbine fans and a full-coverage vapor chamber sustain 140 W at about 32 dB — full performance, near silence. 140W TDP vapor chamber ~32 dB
## By the numbers
04 — By the numbers By the numbers 126 TFLOPS platform AI 128 GB LPDDR5X-8000 16 TB NVMe · PCIe 4.0 140 W TDP, ~32 dB
## Pick the silicon. We'll run it.
Let's build Pick the silicon. We'll run it. Turnkey Edge-AI — fixed time, fixed cost, full responsibility. Talk to our engineers → See all solutions
# Overview · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/index.html
## Agentic-Native SDLC for Regulated Medical Device Engineering #
Overview · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate Agentic-Native SDLC for Regulated Medical Device Engineering # Figure A — Reference Architecture (seven planes) · open SVG A reference framework for transitioning a 1000+ developer, Kubernetes-native medical device organization (GE HealthCare / Siemens Healthineers class) from human-authored, AI-assisted development to a validated, agent-native software development lifecycle — under strict FDA / IEC 62304 obligations, with self-hosted fine-tuned models only , deterministic evaluation, and disciplined GPU/token economics. This is not a vibe-coding playbook. The thesis throughout: generation is cheap; correctness, validation, traceability, and cost control are the engineering. Agents propose; deterministic verifiers and qualified humans dispose. The problem in one paragraph # The organization wants the velocity of agentic development but operates under constraints that rule out the default industry playbook: regulated software demands ≥99.9% release-gate correctness and full auditability; SaaS LLM APIs (Claude/OpenAI/Gemini) are excluded on cost and data-sovereignty grounds, so all inference and training are self-hosted, fine-tuned, open-weight, multi-model ; agentic loops are GPU-expensive , so cost must be engineered down to cost-per-verified-task ; and everything must map cleanly onto IEC 62304, ISO 13485 / FDA QMSR, ISO 14971, FDA CSA, GAMP 5, and 21 CFR Part 11 . The answer in one paragraph # A six-level maturity model (ASMM-Med) moves the org from ungoverned shadow AI → governed assistance → spec-driven bounded automation → orchestrated agentic workflows → validated autonomous agents → a self-optimizing agentic enterprise. Capability is gated by assurance: autonomy can never outrun governance, evaluation, and security. A K8s-native reference architecture serves a tiered fleet of self-hosted fine-tuned models behind a routing gateway, wraps every probabilistic generation in deterministic verifiers + HITL to earn the 99.9% gate, runs agents in zero-trust sandboxes under a policy server , and meters GPU/token cost as a first-class SLO . Document map # # Document Read it for 00 this README.md Executive overview, navigation 01 Requirements Functional, non-functional, regulatory, data, model, and cost requirements (the "shall" statements) 02 Maturity Model (ASMM-Med) The centerpiece: 6 levels × 8 dimensions, gate rules, scoring, KPIs, anti-patterns 03 Reference Architecture K8s-native platform: serving, orchestration, data/RAG, control planes, topology 04 Model Strategy & Fine-Tuning The multi-model fleet, continued-pretrain → SFT → preference → LoRA, reproducibility 05 Evaluation & Validation How 99.9% is earned: deterministic verifiers, eval suites, the assurance argument 06 Agentic Workflows Concrete agent patterns mapped to the SDLC and IEC 62304 activities 07 Security & Compliance Zero-trust, supply chain, prompt-injection defense, CSA/Part 11, autonomy authorization 08 Token & GPU Economics FinOps: routing, caching, quantization, cost-per-green-PR, build-vs-buy math 09 Adoption Roadmap Phased plan, owners, exit criteria, org design, risks Suggested reading order: 02 (frame) → 01 (obligations) → 03/04 (build) → 05 (assurance) → 06 (operation) → 07 (control) → 08 (cost) → 09 (sequence). Seven invariant principles (carried across every document) # **99.9% is a system property, not a model property** — earned at the gate via Generate → Verify → Repair → Gate, not assumed at generation. Determinism wraps probabilism — every check that can be deterministic must be, and on the critical path to merge. Risk-proportional autonomy — IEC 62304 safety class (A/B/C) sets the leash; Class C is always dual human control. Everything an agent does is evidence — immutable, attributable, replayable (21 CFR Part 11 grade). The harness is the product — Agent = Model + Harness; ~90% of behavior and ~100% of assurance live in the harness. Cost is per verified task — the governing metric is cost-per-green-PR , not cost-per-token. Self-hosted, sovereign, reproducible — all models/datasets/training versioned, signed, and regenerable for audit. The 99.9% question, answered up front # No self-hosted open-weight model deterministically produces 99.9%-correct regulated code. We do not try to make it. Instead: Figure B — The 99.9% Release-Gate Assurance Pipeline · open SVG The model is the least trusted component. Trust is manufactured by everything around it. Full treatment in 05-evaluation-and-validation . The cost question, answered up front # Self-hosting trades API OpEx for a GPU fleet (CapEx) + operations (OpEx). We make it pay by: Tiered routing — a 1–8B "reflex" model handles the majority of low-complexity calls; 70B+/MoE "reasoners" are invoked sparingly (see 04 , 08 ). Caching — KV-cache reuse, prompt/semantic caching, retrieval caching. Efficiency — quantization (FP8/INT8/AWQ), speculative decoding, continuous batching, MIG partitioning, scale-to-zero for spiky workloads. Budget guardrails in-loop — hard token/GPU stops per task; eval-cost budgeting; reasoning-effort caps. The right metric — optimize cost-per-green-PR , because an expensive change that passes all gates beats a cheap one that escapes a defect into a regulated product. Scope & assumptions (challenge these) # In scope: AI agents that build/test/document/maintain regulated software (the production/quality-system tooling track). Adjacent (enabled, not detailed): AI shipped inside the device (SaMD) — a separate submission track that reuses the same eval/reproducibility/PCCP muscles ( 07 §"Two regulated tracks"). Platform assumption: existing Kubernetes estate with GPU capacity (on-prem and/or sovereign VPC), service mesh, and a mature CI/CD + QMS. Model assumption: open-weight bases (Qwen / Llama / DeepSeek / Mistral / StarCoder families + a vision-language tier), fine-tuned in-house; no external inference. Regulatory context: US FDA-centric with EU MDR / AI Act awareness; dates as of May–June 2026 (QMSR in effect; FDA CSA final; FDA AI-lifecycle + PCCP guidance available). Authored as an internal engineering/quality reference. Every quantitative threshold (e.g., specific coverage %, GPU counts, SLOs) is a placeholder to be set by the organization's risk and capacity analysis, not a vendor claim. Next → Executive Brief On this page The problem in one paragraph The answer in one paragraph Document map Seven invariant principles (carried across every document) The 99.9% question, answered up front The cost question, answered up front Scope & assumptions (challenge these)
# Executive Brief · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/exec-brief.html
## Executive Brief #
Executive Brief · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate Executive Brief # Board-level summary · one page We will adopt an agent-native software development lifecycle for our regulated medical-device software — capturing the productivity of AI while meeting our FDA / IEC 62304 obligations, keeping all models and data inside our own infrastructure, and controlling cost. This is an engineering and quality discipline, not “vibe coding.” The bet Generation is becoming free; correctness, validation, traceability, and cost control are the engineering . Our advantage comes from the system around the model, not the model alone. Why now AI now writes a large share of new code industry-wide, compressing implementation from weeks to hours. Our competitors are moving. The risk is not adopting too slowly — it is adopting without assurance , which in a regulated context means recalls, findings, and IP leakage. A disciplined framework lets us move fast and stay defensible. What we are building A self-hosted model fleet Open-weight, fine-tuned models served on our existing Kubernetes/GPU platform. No Claude / OpenAI / Gemini APIs — for cost at scale, data sovereignty, and regulatory control. A six-level maturity model ASMM-Med moves us from ungoverned “shadow AI” to validated autonomous agents , in assurance-gated steps each signed off by Engineering, Quality/Regulatory, and Security. Deterministic assurance Every AI output passes deterministic verifiers (compilers, tests, static analysis, formal checks) plus human review before it can ship. The model proposes; the system disposes. Cost as a first-class metric We optimize cost-per-verified-change , not cost-per-token — via tiered model routing, caching, quantization, and in-loop budget limits. The three questions the board will ask Question Our answer How do we hit 99.9% accuracy if the AI is probabilistic? 99.9% is a property of the system , not the model. We wrap every generation in deterministic checks and human checkpoints ( Generate → Verify → Repair → Gate → Sign ). The model is the least-trusted component; trust is manufactured around it. The highest-risk software (IEC 62304 Class C) always requires dual human sign-off. Can we afford the GPU cost? Self-hosting converts per-token vendor billing into an owned, amortizable GPU fleet. A small “reflex” model handles the majority of calls cheaply; large reasoning models are used sparingly. At our scale this is materially cheaper than API pricing — and sovereignty/IP control make it non-optional regardless. Will regulators accept it? Yes, when the agent is treated as validated production software under FDA Computer Software Assurance, with documented intended use, risk-based evidence, full traceability, and immutable 21 CFR Part 11 records. The framework is built to produce that evidence automatically. What it unlocks Throughput — faster implementation, test generation, documentation, and safe modernization of legacy code that was previously “too risky to touch.” Quality — more comprehensive automated test and evaluation coverage than humans can produce in the same time, lowering escape-rate into regulated products. Leverage — smaller teams tackle larger problems; engineers shift from writing code to designing, validating, and directing the systems that produce it. Investment & timeline (illustrative) Horizon Maturity target Focus Quarters 1–2 L1 Governed Assistance Kill shadow AI; stand up self-hosted serving + full logging Quarters 3–4 L2 Spec-Driven Specs in-repo; deterministic evaluation harness in CI Quarters 5–7 L3 Orchestrated Sandboxed agents, model fleet + routing, policy server Quarters 8–11 L4 Validated Autonomous CSA-validate agents; 99.9% gates; full traceability Quarter 12+ L5 Self-Optimizing Closed-loop fine-tuning; cost-optimized at scale Primary investment: GPU capacity, a platform/MLOps team, evaluation engineering, and quality/regulatory integration. Resourcing and figures are placeholders for the funded business case. What we ask of the board Endorse the self-hosted, assurance-gated strategy (no external LLM APIs). Fund Phase 1 (L1) — serving platform, logging, and the end of shadow AI. Affirm the governance principle that autonomy never outruns assurance : no maturity level is granted without joint Engineering + Quality/Regulatory + Security sign-off. Bottom line Structure scales; vibes don’t. AI amplifies our engineering and quality culture — both its strengths and its weaknesses. This framework ensures it amplifies the right ones. Full detail: see the Maturity Model and the document set . ← Previous Overview Next → 01 · Requirements
# 01 · Requirements · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/01-requirements.html
## 01 — Requirements (Normative "Shall" Document) #
01 · Requirements · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate 01 — Requirements (Normative "Shall" Document) # Project: Agentic-Native SDLC for Regulated Medical Device Engineering Status: Baseline v1.0 · Date context: May 2026 Classification: Internal Engineering / Quality Reference Related docs: 02-maturity-model.md · 03-reference-architecture.md · 04-model-strategy-and-finetuning.md · 05-evaluation-and-validation.md · 06-agentic-workflows.md · 07-security-and-compliance.md · 08-token-and-gpu-economics.md · 09-adoption-roadmap.md 1. Purpose, Scope, and How to Read Requirement IDs # 1.1 Purpose # This document is the normative requirements baseline for an agentic-native software development lifecycle (SDLC) serving a 1000+ developer medical-device engineering organization. It defines, in testable "shall" form, what the platform, its model fleet, and its agents must do. It is the contract against which the architecture ( 03-reference-architecture.md ), evaluation system ( 05-evaluation-and-validation.md ), and compliance posture ( 07-security-and-compliance.md ) are judged. 1.2 Scope # In scope Out of scope AI-assisted and AI-autonomous activities across the IEC 62304 software lifecycle (requirements → design → code → test → docs → review → maintenance) Hardware design, electrical/mechanical CAD outside software-controlled subsystems Self-hosted open-weight model fleet, serving, fine-tuning, evaluation, and orchestration Procurement of SaaS LLM services (explicitly prohibited — see §11) Governance, traceability, audit, and validation of the agentic tooling itself (CSA / GAMP 5) Clinical trial design, regulatory submission authoring beyond software evidence Security, observability, and FinOps for the agent platform General IT, HR, or non-engineering enterprise systems 1.3 Requirement ID scheme # Each requirement has the form PREFIX-NNN , a MoSCoW priority, and a mapping to one ASMM-Med maturity level (L0–L5) and one or more of the eight dimensions (D1–D8). Prefix Domain Primary owner FR Functional — what agents do across the SDLC Eng + QA/RA NFR Non-functional — performance, scale, determinism MLOps/Platform REG Regulatory & quality QA/RA DATA Data & knowledge governance MLOps + Security MODEL Model fleet requirements MLOps EVAL Evaluation & assurance QA/RA + Eng SEC Security & zero-trust Security COST FinOps & cost guardrails Finance + Platform OPS Observability & operations Platform MoSCoW priority: M = Must (release-blocking), S = Should, C = Could, W = Won't (this baseline). Conventions: "shall" = mandatory; "should" = recommended; numeric thresholds are org-set placeholders , not vendor claims, and are owned by the named accountable function. Every requirement is verifiable by inspection, demonstration, test, or analysis (method noted in §12). Governing principles (referenced throughout): P1 99.9% is a system property (Generate→Verify→Repair→Gate); P2 determinism wraps probabilism; P3 risk-proportional autonomy by IEC 62304 class A/B/C; P4 everything an agent does is evidence; P5 the harness is the product; P6 cost per verified task; P7 self-hosted, sovereign, reproducible. 2. Stakeholders & Concerns # Stakeholder Primary concerns Key requirement families Veto authority Engineering (Dev) Productivity, low latency, correct suggestions, low friction, not babysitting bad output FR, NFR, OPS No QA / Regulatory Affairs (RA) IEC 62304 conformance, traceability, validated tools, audit-ready records, escape rate REG, EVAL, FR Yes (release gate) Security Zero-trust, no data egress, supply-chain integrity, prompt-injection defense, secrets SEC, DATA, MODEL Yes (deploy gate) MLOps / Platform Reproducibility, model lifecycle, serving SLOs, multi-LoRA, GPU efficiency MODEL, NFR, OPS, DATA No Clinical / Product Safety-class correctness, requirement intent fidelity, time-to-market FR, REG, EVAL Yes (intent) Finance Cost-per-green-PR, GPU capex/opex, budget predictability COST, OPS Yes (budget) A requirement that any veto-holding stakeholder rejects cannot be marked "Accepted" in §12. 3. Functional Requirements (FR) # All FRs are bounded by IEC 62304 safety class (A/B/C) and a defined review posture per P3. "Dual human control" = two qualified humans (author-reviewer separation) for Class C. ID Requirement (shall) Safety-class bounding Review posture MoSCoW ASMM-Med Dim FR-001 Agents shall assist requirements analysis : decompose, classify, detect ambiguity/conflict, and propose acceptance criteria from source specs (incl. PDF/diagram via Tier-V). A/B/C: proposal only Human approves all generated/edited requirements M L2 D3,D5 FR-002 Agents shall generate design support artifacts (interface specs, sequence/architecture sketches, design-decision rationale) traceable to requirements. A/B: draft; C: draft + dual review Human-of-record signs design M L3 D3,D5 FR-003 Agents shall perform code generation scoped to a spec/work item, emitting diffs, not silent edits. A: auto-PR allowed; B: PR + 1 review; C: PR + dual review, no autonomous merge Per class M L2→L4 D5 FR-004 Agents shall perform test generation (unit/integration/property/boundary) mapped to requirements and risk controls (ISO 14971). All classes: tests are evidence, human-confirmed coverage intent Reviewer confirms adequacy M L2 D4,D5 FR-005 Agents shall generate documentation (design history, SDS, API docs, traceability narratives) from code+spec, marked AI-authored. A/B/C: draft QA/RA approves controlled docs M L2 D3 FR-006 Agents shall perform code review producing findings with severity, location, and rationale; review output is advisory, not a gate by itself (P1). All classes Augments, never replaces, human reviewer M L3 D4,D5 FR-007 Agents shall perform refactoring with behavior-preservation evidence (test pass, diff semantics) attached. A: auto; B: review; C: dual review Per class S L3 D5 FR-008 Agents shall perform code/dependency/platform migration with before/after equivalence evidence and rollback plan. B/C: human-gated cutover Migration plan signed by lead S L3 D5 FR-009 Agents shall generate and maintain traceability links (requirement↔design↔code↔test↔risk) and flag gaps. All classes QA/RA owns final trace matrix M L3 D1,D3,D4 FR-010 Every agent action shall produce a Generate→Verify→Repair→Gate record; an action with no verifier shall not pass the gate (P1). All classes System-enforced M L2 D4 FR-011 Agents shall abstain ("I cannot safely complete this") and escalate when confidence/coverage thresholds are unmet, rather than emit low-assurance output. All classes Escalation routed to human M L1 D4,D5 FR-012 Agents shall be orchestrated via the MCP tool plane and A2A for multi-agent workflows with declared, least-privilege tool scopes. All classes Policy-bounded M L3 D5 FR-013 Agents shall produce risk-analysis support (hazard identification candidates, traceable to ISO 14971), human-confirmed. A/B/C: proposal only Risk owner confirms S L3 D1 FR-014 The system shall support human-in-the-loop interrupt/override at any step, with reason captured. All classes Always available M L1 D5,D8 FR-015 Agents shall route tasks across the tiered fleet (Reflex/Worker/Reasoner/Multimodal/Embedding) by task class and cost (P6). All classes System-enforced S L3 D2,D5 4. Non-Functional Requirements (NFR) # ID Requirement (shall) Target (org-set placeholder) MoSCoW ASMM-Med Dim NFR-001 Inline/autocomplete (Tier-S) latency shall meet p95 budget. p95 ≤ 300 ms M L1 D2,D7 NFR-002 Interactive agent step (Tier-M) first-token latency shall meet p95 budget. p95 ≤ 2 s M L2 D2 NFR-003 Reasoning/planning task (Tier-L) end-to-end latency shall meet budget for batch-acceptable workloads. p95 ≤ 60 s S L3 D2 NFR-004 Serving plane shall sustain org-wide concurrent throughput at peak. ≥ 1000 concurrent dev sessions M L2 D2,D7 NFR-005 Control/gate-path availability shall meet SLO. ≥ 99.9% monthly M L2 D7 NFR-006 Gate evaluation shall be deterministic and reproducible : identical inputs + pinned model/LoRA/seed/config → identical gate verdict (P2). 100% verdict reproducibility M L4 D4 NFR-007 Any generated artifact shall be reproducible from recorded {model digest, LoRA, prompt, context snapshot, params, seed} (P7). 100% replayable M L4 D2,D4 NFR-008 Platform shall scale to 1000+ developers via K8s horizontal scaling and KEDA autoscale; idle model pools shall scale to zero. Linear cost-to-load to defined ceiling M L2 D2,D7 NFR-009 Multi-LoRA hot-swap shall serve N task-specialized adapters per base without per-adapter cold redeploy. ≥ defined adapters/base online S L3 D2 NFR-010 Probabilistic model calls shall be wrapped by deterministic harness logic (validators, parsers, policy) so non-determinism cannot reach a gate verdict (P2). No stochastic path to verdict M L2 D4,D5 NFR-011 Recovery: on serving node/GPU failure, in-flight tasks shall be re-queued without evidence loss. RTO ≤ defined; zero record loss S L2 D2,D7 NFR-012 The harness (not just the model) shall be versioned and treated as the product unit (P5); harness changes shall be release-controlled. 100% harness versioned M L3 D5 5. Regulatory & Quality Requirements (REG) # ID Requirement (shall) Regulatory anchor MoSCoW ASMM-Med Dim REG-001 The platform shall enforce the IEC 62304 software safety classification (A/B/C) as a first-class attribute gating autonomy (P3). IEC 62304 M L2 D1 REG-002 All AI-assisted lifecycle activities shall be validated under a risk-based Computer Software Assurance (CSA) approach proportional to intended use and risk. FDA CSA, GAMP 5 (2nd ed) M L2 D1,D4 REG-003 Every agent action shall be recorded as 21 CFR Part 11-grade evidence : attributable, immutable, time-stamped, and replayable (P4). 21 CFR Part 11 M L2 D1,D6 REG-004 The platform shall maintain end-to-end traceability (user need → requirement → design → code → test → risk control) and surface coverage gaps. IEC 62304, ISO 13485/QMSR M L3 D1,D3 REG-005 Each AI tool used in the lifecycle shall be subject to tool validation / qualification with documented intended use, acceptance, and re-validation triggers. CSA, GAMP 5, 21 CFR 820 (QMSR, eff. Feb 2026) M L2 D1,D4 REG-006 Risk management activities shall integrate ISO 14971; AI-proposed hazards/controls shall be human-confirmed before becoming controlled records. ISO 14971 M L3 D1 REG-007 The AI management system governing the fleet shall conform to an AI management system standard . ISO/IEC 42001 S L3 D1 REG-008 AI-enabled-device change management shall support a Predetermined Change Control Plan (PCCP) where models influence device behavior. FDA AI-enabled device guidance + PCCP S L4 D1 REG-009 The platform shall meet applicable EU MDR / EU AI Act obligations for high-risk AI used in device engineering. EU MDR, EU AI Act S L3 D1 REG-010 All controlled records shall have defined retention, version, and signature controls under the QMS. ISO 13485 / FDA QMSR M L2 D1 REG-011 Human accountability shall be preserved: a named qualified human of record shall sign every controlled output; AI is never the signer. 21 CFR Part 11, IEC 62304 M L1 D1,D8 6. Data & Knowledge Requirements (DATA) # ID Requirement (shall) MoSCoW ASMM-Med Dim DATA-001 Training, fine-tuning, and RAG corpora shall be governed : cataloged, licensed, owned, and approved before use. M L2 D3 DATA-002 The platform shall enforce no external egress of source, specs, or model traffic; all inference, training, and storage are self-hosted (P7). M L1 D3,D6 DATA-003 PII/PHI shall be detected (Tier-S redactor) and excluded/masked from corpora, prompts, logs, and evidence stores unless explicitly authorized and controlled. M L1 D3,D6 DATA-004 Every datum used by an agent shall carry provenance (source, version, hash, retrieval timestamp) recorded in the action evidence. M L2 D3,D4 DATA-005 Knowledge bases shall be versioned and snapshot-able so a retrieval context is reproducible for replay (links NFR-007). M L3 D3 DATA-006 Corpus and evidence retention shall follow QMS record-retention policy; deletion shall be controlled and logged. M L2 D1,D3 DATA-007 Corpus quality shall be monitored for drift, staleness, and poisoning ; suspect sources shall be quarantined. S L4 D3,D6 DATA-008 Embeddings/reranking (Tier-E) indices shall be access-controlled per project and safety class. S L3 D3,D6 7. Model Requirements (MODEL) # ID Requirement (shall) MoSCoW ASMM-Med Dim MODEL-001 Only self-hosted open-weight models shall be used; no SaaS LLM API (Claude/OpenAI/Gemini or equivalent) in any lifecycle path (P7, §11). M L1 D2 MODEL-002 All fleet models shall be fine-tunable in-house (full or PEFT/LoRA) on governed corpora. M L2 D2 MODEL-003 Every model and adapter shall be signed and version-pinned (cosign), registered in MLflow with an immutable digest. M L2 D2,D6 MODEL-004 Serving shall support multi-LoRA hot-swap across task-specialized adapters on shared base weights (vLLM/Triton+TensorRT-LLM/KServe). M L3 D2 MODEL-005 Models shall support calibrated abstention — emitting a refusal/low-confidence signal the harness can act on (links FR-011). M L2 D2,D4 MODEL-006 The fleet shall be tiered (Tier-S Reflex 1–8B, Tier-M Worker 14–34B, Tier-L Reasoner 70B+/MoE, Tier-V Multimodal, Tier-E Embedding/Rerank) with documented task→tier routing. M L3 D2,D5 MODEL-007 Multimodal capability (Tier-V) shall ingest diagrams, imaging, PDF specs, and UI for FR-001/FR-002. S L3 D2,D3 MODEL-008 Each model version shall pass acceptance evaluation before promotion to a serving channel (links EVAL-001). M L4 D2,D4 MODEL-009 Model lineage (base → fine-tune dataset → adapter → deployed digest) shall be fully reproducible and recorded (P7). M L4 D2 MODEL-010 Quantization/optimization (TensorRT-LLM) shall not degrade a model below its gated acceptance thresholds without re-validation. S L4 D2,D4 8. Evaluation & Assurance Requirements (EVAL) # ID Requirement (shall) Target (org-set placeholder) MoSCoW ASMM-Med Dim EVAL-001 Release gates shall be deterministic and produce a binary, reproducible verdict (P1, P2). 100% reproducible M L4 D4 EVAL-002 The system release-gate acceptance correctness shall meet the org threshold as a system property via Generate→Verify→Repair→Gate, not from any single model (P1). ≥ 99.9% M L4 D4 EVAL-003 Escape rate (defects passing the gate into controlled artifacts) shall be measured and bounded. ≤ org-set ceiling M L4 D4 EVAL-004 Golden datasets per task/safety-class shall exist, be versioned, and gate model/harness promotion. 100% coverage of gated tasks M L4 D4 EVAL-005 An LLM shall never be the sole gate ; gates shall combine deterministic verifiers (build/test/static analysis/policy) with optional model judgment as advisory only (§11). Enforced M L2 D4 EVAL-006 Gate verifiers shall include compile/build, test execution, static analysis, and policy (OPA/Gatekeeper) checks. All present M L3 D1,D4 EVAL-007 Evaluation results shall be evidence (P4): stored immutable, attributable to model/harness digests, replayable. 100% M L4 D4 EVAL-008 Continuous evaluation shall detect model/behavioral drift post-deployment and trigger re-validation. Monitored S L5 D4 EVAL-009 Repair loops shall be bounded (max iterations/budget); on exhaustion the task shall escalate to human (links FR-011, COST). Bounded M L2 D4,D7 9. Security Requirements (SEC) # ID Requirement (shall) Mechanism MoSCoW ASMM-Med Dim SEC-001 The platform shall be zero-trust : every workload identity authenticated and authorized per call. Istio mesh + SPIFFE/SPIRE M L2 D6 SEC-002 Agent code execution shall run in isolated sandboxes with no ambient credentials or network. gVisor/Kata M L2 D6 SEC-003 Supply chain shall be signed and attested : artifacts, models, containers via Sigstore/cosign + SLSA provenance + SBOM. cosign/SLSA/SBOM M L2 D6 SEC-004 The platform shall implement prompt-injection and tool-abuse defenses (input sanitization, tool allow-lists, output schema validation, least privilege). MCP scopes + validators M L3 D5,D6 SEC-005 Secrets shall be managed centrally and never appear in prompts, logs, or evidence. HashiCorp Vault M L1 D6 SEC-006 The audit/evidence store shall be immutable and tamper-evident (append-only, hash-chained) (P4). Append-only + cosign M L2 D1,D6 SEC-007 Policy shall be enforced at admission and runtime via a policy server (OPA/Gatekeeper); no policy bypass path. OPA/Gatekeeper M L2 D6 SEC-008 The platform shall meet medical-device cybersecurity obligations and network segmentation. IEC 62443, FDA §524B S L3 D6 SEC-009 Tool plane (MCP) and multi-agent (A2A) calls shall enforce least-privilege, declared scopes , audited per invocation. MCP/A2A policy M L3 D5,D6 SEC-010 Agents shall never autonomously merge or release Class C changes (links FR-003, §11). Branch policy + gate M L3 D1,D6 10. Observability & Cost Requirements (OPS / COST) # ID Requirement (shall) Target (org-set placeholder) MoSCoW ASMM-Med Dim OPS-001 Every agent action and model call shall emit OpenTelemetry traces correlatable end-to-end (request→tools→model→gate→evidence). 100% traced M L2 D7 OPS-002 The platform shall expose GPU utilization, queue depth, and tokens/sec per tier and per tenant. Dashboards live M L2 D7 OPS-003 Drift, error-rate, abstention-rate, and escape-rate shall be observable in near-real-time. SLO dashboards S L4 D4,D7 COST-001 The platform shall compute cost-per-green-PR (cost per verified task) as the primary efficiency KPI (P6). Reported per team M L3 D7 COST-002 Budget guardrails shall enforce per-team/per-task token & GPU ceilings; overruns throttle or escalate, never silently spend. Hard ceilings M L2 D7 COST-003 Routing shall prefer the cheapest tier that meets the quality gate (links FR-015, MODEL-006). Enforced S L3 D2,D7 COST-004 Idle GPU pools shall scale to zero ; cost attribution shall be tenant-accurate. KEDA + chargeback S L2 D7 COST-005 Repair/retry loops shall be cost-bounded (links EVAL-009); runaway loops shall halt and escalate. Bounded M L2 D4,D7 11. Constraints & Explicit Non-Goals # 11.1 Hard constraints (shall) # ID Constraint CON-001 No SaaS/hosted LLM APIs (Claude, OpenAI, Gemini, or equivalent) in any lifecycle path. Open-weight, self-hosted only. CON-002 No external network egress of source, specs, PHI/PII, or model traffic. CON-003 No LLM as a sole gate : a deterministic verifier set must back every release decision (EVAL-005). CON-004 No autonomous merge or release of Class C software by an agent; Class C requires dual human control (SEC-010, FR-003). CON-005 No agent action without replayable Part 11-grade evidence (REG-003). CON-006 No model/adapter deployment without signing, registry entry, and acceptance evaluation (MODEL-003, MODEL-008). CON-007 No non-deterministic path may reach a gate verdict (NFR-010, P2). 11.2 Explicit non-goals (this baseline) # Non-goal Rationale Fully unattended Class C autonomy Prohibited by P3; revisit only with regulatory precedent. "Vibe coding" / unbounded freeform generation Counter to spec-driven, evidence-bound philosophy. General-purpose chatbot assistant outside the SDLC Out of scope; no validation basis. Vendor-managed model hosting Conflicts with P7 sovereignty. Replacing human reviewers/signers AI augments; humans remain accountable (REG-011). 12. Acceptance Criteria Summary # Verification methods: I = Inspection, D = Demonstration, T = Test, A = Analysis. Requirement set Acceptance criterion (must pass for baseline sign-off) Method Priority FR-001…015 Each SDLC activity demonstrated with safety-class bounding and correct review posture; abstention and human override exercised. D, T M NFR-001…012 Latency/throughput/availability SLOs met under load test; gate verdict + artifact reproducibility shown bit-stable on replay. T, A M REG-001…011 Traceability matrix complete with no orphan links; CSA/tool-validation dossiers present; Part 11 evidence replayed; human-of-record signatures verified. I, A M DATA-001…008 No-egress proven by network policy test; PHI/PII redaction validated; provenance present on sampled actions; corpus versioning replayable. T, I M MODEL-001…010 Fleet confirmed open-weight/self-hosted; signatures and registry digests verified; multi-LoRA hot-swap demonstrated; abstention signal observed; lineage reproducible. I, D, T M EVAL-001…009 Gate determinism proven (identical inputs → identical verdict); system acceptance ≥ 99.9% on golden sets; escape rate within ceiling; no LLM-sole-gate path exists. T, A M SEC-001…010 Zero-trust identity enforced; sandbox isolation tested; SLSA/SBOM/cosign present; prompt-injection suite passed; Class C autonomous-merge attempt blocked. T, I M OPS/COST-001…005 End-to-end traces present; cost-per-green-PR reported; budget guardrail throttle demonstrated; scale-to-zero and chargeback verified. D, T M CON-001…007 Each hard constraint shown enforced (negative tests: egress blocked, SaaS call blocked, LLM-sole-gate rejected, Class C auto-merge rejected). T M Baseline sign-off requires all Must requirements Accepted with no open veto from any §2 veto-holder (QA/RA, Security, Finance, Clinical/Product), and traceability of every requirement to at least one verification record per P4. End of 01-requirements.md — proceed to 02-maturity-model.md for the ASMM-Med level definitions that scope phased rollout in 09-adoption-roadmap.md . ← Previous Executive Brief Next → 02 · Maturity Model On this page 1. Purpose, Scope, and How to Read Requirement IDs 2. Stakeholders & Concerns 3. Functional Requirements (FR) 4. Non-Functional Requirements (NFR) 5. Regulatory & Quality Requirements (REG) 6. Data & Knowledge Requirements (DATA) 7. Model Requirements (MODEL) 8. Evaluation & Assurance Requirements (EVAL) 9. Security Requirements (SEC) 10. Observability & Cost Requirements (OPS / COST) 11. Constraints & Explicit Non-Goals 12. Acceptance Criteria Summary
# 02 · Maturity Model · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/02-maturity-model.html
## ASMM-Med — Agentic SDLC Maturity Model for Regulated Medical Device Engineering #
02 · Maturity Model · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate ASMM-Med — Agentic SDLC Maturity Model for Regulated Medical Device Engineering # Figure D — ASMM-Med Maturity Staircase · open SVG Audience: Engineering leadership, Quality/Regulatory (QA/RA), MLOps/Platform, and Security at a 1000+ developer medical-device organization (GE HealthCare / Siemens Healthineers class). Scope: AI agents used to build, test, document, and maintain regulated software (the tooling side of the SDLC), running on a Kubernetes cloud-native platform with self-hosted, fine-tuned open-weight models only (no Claude/OpenAI/Gemini SaaS APIs). Companion docs: 01-requirements · 03-reference-architecture · 04-model-strategy-and-finetuning · 05-evaluation-and-validation · 06-agentic-workflows · 07-security-and-compliance · 08-token-and-gpu-economics · 09-adoption-roadmap 1. Purpose # This maturity model gives a regulated medical-device engineering organization a defensible, auditable path from ad-hoc AI assistance to a validated, agent-native software development lifecycle. It is explicitly not a vibe-coding model . Every level raises both capability and assurance, because in this domain unverified velocity is a liability, not an asset. The model answers four questions leadership repeatedly asks: How do we get to ≥99.9% release-gate correctness when the underlying models are probabilistic? How do we stay inside FDA / IEC 62304 / ISO 13485 obligations while letting agents touch regulated code? How do we control GPU/token cost when agentic loops are expensive and we self-host? What does "good" look like at each step , so we can fund, audit, and de-risk the transition? 2. Foundational design principles # These principles are invariant across all levels and bind the rest of the documentation set. P1 — 99.9% is a system property , not a model property # No single open-weight model will deterministically hit 99.9% functional correctness on regulated code. The target is met by the system : a probabilistic generator wrapped in deterministic verifiers and human checkpoints . Figure F — How Released Correctness Is Earned · open SVG The model's job is to propose ; the harness's job is to dispose . Correctness is earned at the gate, not assumed at generation. See 05-evaluation-and-validation for the full assurance argument. P2 — Determinism wraps probabilism # Wherever a check can be deterministic (type systems, unit/property/mutation tests, static analysis, schema validation, policy-as-code, formal/specification checks), it must be, and it sits on the critical path to merge. Probabilistic judgment (LLM-as-judge) is permitted only as a secondary, escalating signal, never as a sole gate for risk-classified changes. P3 — Risk-proportional autonomy (IEC 62304 safety class drives the leash length) # Agent autonomy is a function of the safety class of the artifact being modified: Class A (no injury possible): high autonomy permitted at higher maturity. Class B (non-serious injury): bounded autonomy, mandatory human review. Class C (death / serious injury): agent may propose and evidence , but a qualified human always authors the merge decision; dual control required. P4 — Everything an agent does is evidence # Every prompt, context bundle, model+adapter version, tool call, verifier result, and human decision is captured as immutable, attributable, replayable record (21 CFR Part 11-grade). If it isn't logged, it didn't happen — and it can't ship in a regulated product. P5 — The harness is the product (Agent = Model + Harness) # ~90% of behavior and ~100% of assurance comes from the harness (instructions, tools, sandboxes, policies, evals, observability), not the raw model. Investment, validation, and change control concentrate on the harness. Models are hot-swappable inputs; the harness is the controlled system. P6 — Cost is measured per resolved, verified task , not per token # Self-hosting converts OpEx (API calls) into CapEx+OpEx (GPU fleet). The governing metric is cost-per-green-PR (a change that passes all deterministic gates and human review), not cost-per-token. Routing, caching, and tiering exist to minimize that. See 08-token-and-gpu-economics . P7 — Self-hosted, sovereign, reproducible # All inference and training run inside the organization's trust boundary (on-prem or sovereign VPC). Models, datasets, and training runs are versioned, signed, and reproducible so that any artifact an agent produced can be regenerated and defended in an audit or a recall investigation. 3. The six maturity levels # Level Name One-line essence Dominant operating mode Autonomy ceiling L0 Ad-hoc Assistance Ungoverned, shadow AI Individual, in-editor Suggestions only L1 Governed Assistance Sanctioned self-hosted assist Conductor (human types) Inline completion L2 Spec-Driven Bounded Automation Specs + single agents on reviewable tasks Conductor + bounded tasks Single-step, full review L3 Orchestrated Agentic Workflows Sandboxed multi-step agents open PRs Orchestrator Multi-step, HITL gates L4 Validated Autonomous Agents Agents validated as CSA software tools Orchestrator + validated autonomy Autonomous within validated bounds L5 Self-Optimizing Agentic Enterprise Closed-loop, eval-driven, cost-optimal Fleet governance Self-improving under PCCP-style control L0 — Ad-hoc Assistance (the starting reality, not a goal) # Developers use whatever in-IDE autocomplete they can reach. No central policy, no logging, no model governance, and likely shadow use of public SaaS endpoints — an IP-leakage and compliance incident waiting to happen. There is no audit trail tying generated code to a model version. This level is non-compliant by default and the transformation's first job is to extinguish it. L1 — Governed Assistance # A sanctioned, self-hosted code-assistance model (Tier-S/Tier-M, see 04 ) is offered through the IDE behind SSO, with prompt/response logging, an Acceptable-Use Policy, and DLP/PII guards. Humans still author 100% of production code; the AI accelerates typing and lookup. The win is eliminating shadow AI and establishing the logging substrate that all later assurance depends on. L2 — Spec-Driven Bounded Automation # The organization adopts Spec-Driven Development : specs/ , AGENTS.md /rule files, and BDD/Gherkin acceptance criteria live in the repo as the source of truth. Single-step agents perform bounded, individually reviewable tasks — test generation, documentation, mechanical refactors, boilerplate — and a deterministic evaluation harness runs in CI . Every agent output is a normal PR under full human review. This is the first level where agents write code that can ship, and the first level where the eval harness exists. L3 — Orchestrated Agentic Workflows # Agents run multi-step, in sandboxes (ephemeral, egress-restricted), call tools through an MCP tool plane , and are routed across a fleet of fine-tuned models by task complexity. A policy server gates every tool call (structural + semantic). Agents open PRs autonomously but human-in-the-loop checkpoints are mandatory at risk-defined boundaries. Trajectory + output evaluations join the deterministic gates. The org now ships agent-produced changes at scale with managed risk. L4 — Validated Autonomous Agents (the regulated inflection point) # Each production agent is treated as a software tool used in production of a medical device and is validated under FDA Computer Software Assurance (CSA) : documented intended use , risk-based test evidence, and recorded validation. Agents operate autonomously within validated bounds for their safety-class envelope, coordinate via A2A multi-agent patterns, and are continuously gated by ≥99.9% eval thresholds with full IEC 62304 traceability (requirement → spec → code → test → eval → release). Class C work still requires dual human control. This is where "agentic" becomes "regulated-grade." L5 — Self-Optimizing Agentic Enterprise # A closed loop connects production telemetry → eval failures → curated fine-tuning data → candidate adapters → automated eval-driven promotion — all under a Predetermined Change Control Plan (PCCP)-style governance so model/harness evolution is pre-authorized and auditable. The fleet self-tunes for quality and cost-per-green-PR , observability is enterprise-wide, and agents participate in their own improvement under human governance. Maturity here is operational discipline at scale , not raw autonomy. 4. The eight capability dimensions # Maturity is assessed independently along eight axes. An organization is rarely uniform; the floor across safety-relevant dimensions (D1, D4, D6) governs what autonomy is actually permitted , regardless of how advanced D2/D5 are. # Dimension What it measures D1 Governance, Quality & Regulatory Compliance QMS integration, IEC 62304/ISO 13485/QMSR alignment, CSA validation of tools, traceability, change control D2 Model Infrastructure & MLOps Self-hosted serving, fine-tuning pipeline, model registry, reproducibility, multi-LoRA, GPU platform D3 Context & Knowledge Engineering Specs, rule files, RAG over code/docs/regulatory corpus, memory, context hygiene D4 Evaluation, Validation & Assurance Deterministic verifiers, eval suites, trajectory eval, acceptance criteria, the 99.9% gate, abstention D5 Agentic Orchestration & Tooling Single→multi-agent, MCP/A2A, sandboxing, HITL design, workflow engine D6 Security & Zero-Trust Identity, egress control, prompt-injection defense, supply chain, secrets, IEC 62443, audit immutability D7 Observability & FinOps Tracing, eval dashboards, token/GPU metering, routing economics, budget guardrails D8 People, Skills & Operating Model Roles (conductor/orchestrator), review culture, training, approval-fatigue controls 5. The maturity matrix (core artifact) # Each cell is the exit criterion for that dimension at that level (you reach the level only when every dimension meets at least that level's descriptor; see §6). D1 — Governance, Quality & Regulatory Compliance # L Descriptor L0 No policy; shadow AI; no link between generated code and a model version. Non-compliant. L1 AUP published; AI-assist logged; QMS acknowledges AI tooling; data-handling/IP policy enforced (no external endpoints). L2 AI-tool use captured in the DHF/quality records; SDD specs are controlled documents; SOP for "AI-assisted change" exists; risk assessment (ISO 14971) covers AI tooling. L3 Risk-based tool classification per GAMP 5; agent actions mapped to IEC 62304 activities; change-control board reviews harness changes; ISO/IEC 42001 AI-management controls adopted. L4 Each agent validated under FDA CSA with documented intended use + evidence; full requirement-to-release traceability; agents recognized in ISO 13485/QMSR as validated production tooling; periodic revalidation triggers defined. L5 PCCP-style predetermined change control governs model/harness evolution; continuous compliance monitoring; automated audit-evidence generation; regulatory-grade reproducibility of any historical agent action. D2 — Model Infrastructure & MLOps # L Descriptor L0 None / external SaaS. L1 Single self-hosted model served on K8s GPU (vLLM/Triton) behind an internal gateway; basic autoscaling. L2 Model registry (MLflow) with versioned, signed models; first fine-tunes (LoRA/SFT) on internal corpus; reproducible serving images. L3 Tiered model fleet (S/M/L/V/E) with multi-LoRA hot-swap ; routing gateway; quantization (FP8/INT8/AWQ); distributed training (Ray/Kueue); model cards mandatory. L4 Reproducible, validated training pipelines (locked datasets/seeds, signed lineage) suitable as validation evidence; canary + shadow deployment; rollback guarantees; per-domain adapters. L5 Continuous fine-tuning from production signal; automated candidate training; eval-gated promotion; fleet-level capacity/cost optimization; data flywheel under governance. D3 — Context & Knowledge Engineering # L Descriptor L0 Whatever is in the editor buffer. L1 Curated system prompts / org coding standards injected; no retrieval. L2 specs/ + AGENTS.md + BDD criteria in-repo; static rule files; basic code RAG. L3 Governed RAG over code, design docs, SOPs, and the regulatory corpus ; static-vs-dynamic context split; Agent Skills with progressive disclosure; PII/PHI context hygiene middleware. L4 Validated knowledge sources (controlled, versioned, access-scoped); retrieval provenance recorded as evidence; graph-native code understanding for large legacy estates. L5 Self-curating knowledge base; context quality measured and optimized; memory governed and access-audited fleet-wide. D4 — Evaluation, Validation & Assurance # L Descriptor L0 "Looks right." None. L1 Lint + existing CI on human-authored code; no AI-specific eval. L2 Deterministic eval harness in CI (build, unit/property tests, type-check, SAST); per-task acceptance criteria defined; AI changes can't merge red. L3 Output + trajectory evals ; golden datasets with rubrics; mutation testing; LLM-as-judge as secondary signal; abstention/escalation wired in; escape-rate tracked. L4 ≥99.9% release-gate enforced via generate→verify→repair + N-sample self-consistency + verifier voting; eval suites are controlled validation artifacts ; statistical acceptance (CI bounds) per safety class; HITL mandatory for Class B/C. L5 Continuous eval against live failure modes; auto-regression capture; eval-driven model promotion; drift detection closes the loop. D5 — Agentic Orchestration & Tooling # L Descriptor L0 None. L1 Inline completion only; no tool use. L2 Single-step agents , bounded tasks, no autonomous multi-file changes; tools read-only or PR-only. L3 Multi-step agents in sandboxes (gVisor/Kata, ephemeral, egress-deny); MCP tool plane ; policy server gates every call; HITL checkpoints designed in. L4 Multi-agent (A2A) with role decomposition (planner/coder/test/review); validated tool catalog; deterministic hooks at lifecycle points; durable sessions/memory. L5 Self-orchestrating workflows with governed dynamic planning; sub-agent fleets; automated decomposition; coordination cost-optimized. D6 — Security & Zero-Trust # L Descriptor L0 Uncontrolled; IP egress risk. L1 SSO, RBAC, TLS, prompt/response logging, DLP/secret-scanning on inputs. L2 Per-repo scoping; signed images; secret management (Vault); no agent write access to prod. L3 Zero-trust mesh (mTLS, SPIFFE/SPIRE); OPA/Gatekeeper ; sandboxed exec with egress allow-lists; prompt-injection & context-poisoning defenses ; PII/PHI masking; SBOM. L4 Supply-chain assurance (SLSA, Sigstore/cosign signed models+datasets+artifacts); semantic policy gating; IEC 62443 + FDA premarket cybersecurity (§524B) alignment; WORM immutable audit ; dual-control for high-risk tool calls. L5 Continuous adversarial testing (red-team agents); automated anomaly/drift response; tamper-evident, fleet-wide, self-defending posture. D7 — Observability & FinOps # L Descriptor L0 None. L1 Basic request logging + GPU utilization metrics. L2 Per-team usage dashboards; cost attribution; eval pass/fail visibility. L3 End-to-end tracing (OpenTelemetry) of agent trajectories; token/GPU cost per task ; routing telemetry; budget alerts. L4 Cost-per-green-PR as a board metric; per-model/per-tier economics; SLOs on latency/quality/cost; capacity forecasting; budget guardrails enforced in-loop (hard stops). L5 Closed-loop cost optimization; automated routing/quantization decisions; ROI attribution; predictive scaling. D8 — People, Skills & Operating Model # L Descriptor L0 Individual experimentation. L1 Training on sanctioned tool + AUP; champions identified. L2 Spec-writing & review skills built; "conductor" mode normalized; review checklists for AI output. L3 Orchestrator role emerges; ownership split (e.g., API vs UX) to reduce merge conflict; approval-fatigue controls (digital quiet hours, batched review). L4 Formal roles: Agent Steward, Eval Owner, Harness Engineer; no-blame integration culture ; reviewers trained on AI failure modes (hallucinated deps, plausible-wrong logic). L5 Hiring/skills reframed around judgment over implementation ; org-wide agentic literacy; continuous capability development; humans focus on architecture, risk, and verification. 6. Level-gate rules (how you actually "are" at a level) # Floor, not average. Your level on any dimension is its lowest satisfied descriptor. Your overall operating level is the minimum across D1, D4, and D6 (the safety/assurance/security triad). You may invest ahead in D2/D5, but you may not grant autonomy beyond what D1/D4/D6 support. (This is the single most important rule — it prevents capability from outrunning assurance.) Evidence-based promotion. Advancing a level requires a documented assessment (§7) signed by Engineering + QA/RA + Security. No self-attestation. Safety-class gating. Even at L4/L5, autonomy is capped per IEC 62304 class (P3). L5 does not mean "agents merge Class C code unattended" — it never does. Reversibility. Every level must support rollback to the prior level's controls if eval escape-rate or incident metrics breach threshold. 7. Assessment & scoring method # Cadence: quarterly self-assessment; annual independent (internal audit) assessment; event-driven re-assessment after any Sev-1 AI-attributable defect or recall-relevant escape. Instrument: 8 dimensions × 6 levels rubric (this document), scored 0–5 each, with required evidence artifacts per cell (logs, validation records, eval reports, signed model lineage, policy configs). Scoring outputs: Dimension scores (radar chart) → reveals imbalance. Governing level = min(D1, D4, D6) . Capability level = mean(all) (informational only). Autonomy authorization = derived table mapping (governing level × safety class) → permitted agent actions (lives in 07-security-and-compliance ). Gate review: QA/RA holds veto. Security holds veto. Promotion requires both plus Engineering sign-off. Illustrative target trajectory (informational) # Quarter (from program start) Target governing level Q1–Q2 L1 (kill shadow AI, stand up serving + logging) Q3–Q4 L2 (SDD + deterministic eval harness) Q5–Q7 L3 (sandboxed orchestration + fleet + policy server) Q8–Q11 L4 (CSA validation of agents, 99.9% gates) Q12+ L5 (closed-loop, cost-optimized) See 09-adoption-roadmap for the detailed phased plan, owners, and exit criteria. 8. Per-level KPIs # Level Capability KPIs Assurance / cost KPIs L1 % devs on sanctioned tool; shadow-AI incidents → 0 100% prompts logged; IP-egress events = 0 L2 % repos with specs/ + AGENTS.md ; agent-PR acceptance rate CI deterministic-gate coverage ≥ X%; zero red merges L3 % tasks via orchestrated agents; HITL checkpoint adherence trajectory-eval pass rate; escape-rate ; cost-per-task L4 autonomous-task throughput within validated bounds release-gate ≥99.9% ; validation-evidence completeness; revalidation on time L5 fine-tune cycle time; auto-promotion rate cost-per-green-PR trend ↓; drift MTTR; audit-evidence automation % 9. Anti-patterns (explicitly disallowed) # Autonomy ahead of assurance — granting L3+ autonomy while D4/D6 sit at L1/L2. Forbidden by §6.1. LLM-as-sole-gate — using a model to approve a model's output on Class B/C code. Violates P2. Unversioned models — serving a model you cannot reproduce or sign. Violates P7 and CSA. Eval theater — a passing demo presented as a passing eval. Evals need rubrics + statistical acceptance (D4-L4). Context dumping — stuffing 100k-token repos into prompts; burns GPU and degrades accuracy. See 08 . Approval-fatigue reflex-clicking — unbatched micro-approvals leading reviewers to rubber-stamp. Controlled at D8-L3. Shadow SaaS — any call to external LLM endpoints. Hard-blocked at network layer (D6). 10. Regulatory mapping (orientation) # Standard / regulation How this model engages it IEC 62304 (medical device software lifecycle) Agent actions mapped to lifecycle activities; safety class drives autonomy (P3); traceability at D1-L4 ISO 13485 / FDA QMSR (21 CFR 820, effective Feb 2026) AI agents recognized as validated production/quality tooling; SOPs + DHF integration ISO 14971 (risk management) AI-tooling failure modes in the risk file; mitigations = deterministic gates + HITL FDA Computer Software Assurance (CSA) Risk-based validation of each agent as production/QS software (D1-L4) GAMP 5 (2nd ed.) Risk-based, critical-thinking validation approach for the tool category 21 CFR Part 11 Immutable, attributable, signed e-records of all agent actions (P4, D6-L4) ISO/IEC 42001 (AI management system) Org-level AI governance controls (D1-L3+) IEC 62443 / FDA premarket cybersecurity (§524B) Zero-trust + supply-chain assurance for the agent platform (D6-L4) FDA AI-enabled device guidance + PCCP Pattern reused at D1-L5 for governed model evolution of the dev toolchain Critical distinction: this model governs AI that builds the device (production/quality-system software). AI shipped inside the device (SaMD/AI-enabled function) is a separate submission-bearing track — but the same assurance muscles (eval rigor, reproducibility, PCCP) directly enable it. See 07-security-and-compliance §"Two regulated tracks." 11. How to read the rest of the set # What we must satisfy → 01-requirements What we build it on → 03-reference-architecture What models, and how we make them ours → 04-model-strategy-and-finetuning How we earn 99.9% → 05-evaluation-and-validation How agents actually work day-to-day → 06-agentic-workflows How we keep it safe & compliant → 07-security-and-compliance How we afford it → 08-token-and-gpu-economics How we get there → 09-adoption-roadmap ← Previous 01 · Requirements Next → 03 · Reference Architecture On this page 1. Purpose 2. Foundational design principles 3. The six maturity levels 4. The eight capability dimensions 5. The maturity matrix (core artifact) 6. Level-gate rules (how you actually "are" at a level) 7. Assessment & scoring method 8. Per-level KPIs 9. Anti-patterns (explicitly disallowed) 10. Regulatory mapping (orientation) 11. How to read the rest of the set
# 03 · Reference Architecture · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/03-reference-architecture.html
## 03 — Reference Architecture #
03 · Reference Architecture · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate 03 — Reference Architecture # Figure A — Reference Architecture (seven planes) · open SVG Project: Agentic-Native SDLC for Regulated Medical Device Engineering Document: Reference Architecture Status: Controlled — Engineering/Quality Reference Revision date: May 2026 Audience: Platform, MLOps, Security, and Quality engineering leads (1000+ developer org) Sibling documents: 01-requirements.md · 02-maturity-model.md · 04-model-strategy-and-finetuning.md · 05-evaluation-and-validation.md · 06-agentic-workflows.md · 07-security-and-compliance.md · 08-token-and-gpu-economics.md · 09-adoption-roadmap.md 1. Architecture goals, constraints, and the layered view # 1.1 Goals # This reference architecture is the buildable expression of the seven principles defined in 01-requirements.md . It must: Deliver ≥99.9% release-gate correctness as a system property (P1) by composing Generate → Verify → Repair → Gate , not by trusting any single model output. Wrap probabilistic model behavior in deterministic control (P2): every gate, policy decision, and promotion is reproducible from recorded inputs. Enforce risk-proportional autonomy (P3) keyed to IEC 62304 software safety class A/B/C, with Class C always under dual human control . Treat every agent action as Part 11-grade evidence (P4): immutable, attributable, time-stamped, reconstructable. Treat the harness as the product (P5): Agent = Model + Harness . The architecture invests in the harness/control plane, not just the model. Optimize cost per verified task / cost-per-green-PR (P6), making GPU and token spend a first-class, observable quantity (see 08-token-and-gpu-economics.md ). Remain self-hosted, sovereign, and reproducible (P7): open-weight, fine-tuned models only; no external SaaS LLM APIs. 1.2 Hard constraints (non-negotiable) # # Constraint Architectural consequence C1 Open-weight, self-hosted, fine-tuned models only — no Claude/OpenAI/Gemini SaaS All inference runs inside the cluster on the Model Fleet; no egress to LLM providers. Network policy default-deny to public LLM endpoints. C2 ≥99.9% release-gate correctness Multi-stage verification (sandbox + policy + eval) gates every artifact; no model-only merges. C3 GPU/token cost is first-class Tiered model fleet + routing gateway + scale-to-zero + FinOps telemetry on every span. C4 Deterministic evaluation Pinned model digests, fixed seeds, hermetic test environments, content-addressed eval datasets. C5 Sovereign / air-gap capable Every dependency mirrorable; no runtime dependency on internet reachability. 1.3 The seven-plane layered view # flowchart TB subgraph DI["1 · Developer Interface Plane"] IDE["IDE plugins (VS Code / JetBrains)"] CLI["Agent CLI"] WEB["Review & Approval Web UI"] CICD["CI/CD triggers (Argo Events)"] end subgraph AO["2 · Agent / Orchestration Plane"] RUNTIME["Agent Runtime (planner/executor)"] A2A["A2A multi-agent bus"] MCP["MCP Tool Plane"] end subgraph HC["3 · Harness / Control Plane"] ROUTER["Model-Routing Gateway"] POLICY["Policy Server (OPA/Gatekeeper)"] AUTHZ["Autonomy Authorization Service"] EVAL["Eval / Verification Service"] SANDBOX["Sandbox Execution (gVisor/Kata)"] end subgraph MS["4 · Model-Serving Plane"] VLLM["vLLM"] TRITON["Triton + TensorRT-LLM"] KSERVE["KServe + multi-LoRA"] end subgraph DK["5 · Data / Knowledge Plane"] KG["Code Knowledge Graph"] VEC["Vector / Embedding Store"] REG["Regulatory Corpus"] REGISTRY["MLflow Model Registry & Lineage"] end subgraph PI["6 · Platform / Infra Plane"] K8S["Kubernetes + GPU Operator"] RAY["Ray + Kueue"] MESH["Istio + SPIFFE/SPIRE"] VAULT["Vault · KEDA · Argo"] end subgraph GA["7 · Governance / Audit Plane"] WORM["WORM Evidence Store (Part 11)"] OTEL["OpenTelemetry pipeline"] FINOPS["FinOps / cost ledger"] SUPPLY["Sigstore/cosign · SLSA · SBOM"] end DI --> AO --> HC HC --> MS HC --> DK AO --> DK MS --> PI HC --> PI AO -.evidence.-> GA HC -.evidence.-> GA MS -.telemetry.-> GA Plane responsibilities at a glance: Plane Owns Does not own Developer Interface Intent capture, review, approval surfaces Model selection, policy Agent / Orchestration Trajectory planning, tool calls, multi-agent coordination Inference, gating verdicts Harness / Control Routing, policy, autonomy authz, verification, gating Model weights, business logic Model-Serving Inference, batching, LoRA hot-swap Trajectory, gating Data / Knowledge Retrieval, lineage, registry Generation Platform / Infra Scheduling, identity, secrets, scaling Domain semantics Governance / Audit Immutable evidence, telemetry, supply chain Live request handling 2. Logical components per plane # 2.1 Developer Interface Plane # Component Tech Responsibility IDE integration VS Code / JetBrains extensions, LSP bridge Inline intent capture, diff preview, approval prompts, trajectory visualization Agent CLI Self-hosted CLI binary (mTLS to mesh) Headless agent invocation, batch tasks, CI usage Review/Approval Web UI Internal SPA behind Istio + OIDC HITL review, autonomy-class approvals, evidence inspection CI/CD triggers Argo Events + Argo Workflows Event-driven agent runs (PR opened, requirement changed) All interface clients are thin: they hold no model credentials and reach the cluster only through the mesh ingress with SPIFFE-issued identity. 2.2 Agent / Orchestration Plane # Component Tech Responsibility Agent Runtime Custom planner/executor on K8s Job / Pod , Ray actors for fan-out Owns the trajectory: plan → act → observe → repair loop MCP Tool Plane Model Context Protocol servers (one per capability) Typed, permissioned tool surface: repo.read , repo.write , test.run , kg.query , eval.submit , vault.lease A2A bus Agent-to-Agent protocol over the Istio mesh Specialist agents (coder, reviewer, test-author, requirements-tracer) coordinate Tool calls never hit infrastructure directly; they are mediated by MCP servers that enforce per-tool scopes and emit evidence (P4). 2.3 Harness / Control Plane (the product, P5) # Component Tech Responsibility Model-Routing Gateway Custom gateway + classifier (Tier-S model) in front of serving Classify request → select tier/LoRA → enforce cost budget Policy Server OPA/Gatekeeper + dedicated policy server (Rego bundles) Admission of tool calls, autonomy decisions, write permissions Autonomy Authorization Service Custom service keyed to IEC 62304 class Decides allowed autonomy level per task (P3); Class C → dual human Eval / Verification Service Deterministic eval harness (see 05 ) Runs gate suites; emits pass/fail with evidence Sandbox Execution gVisor / Kata Containers, ephemeral namespaces Hermetic build/test/exec of generated artifacts 2.3.1 Model-routing gateway — classification and tier selection # flowchart LR REQ["Agent request
(task + context budget)"] --> CLS{"Classifier
(Tier-S Reflex)"} CLS -->|"trivial / lint / format"| S["Tier-S Reflex
1-8B"] CLS -->|"bounded code edit / unit test"| M["Tier-M Worker
14-34B"] CLS -->|"design / multi-file / reasoning"| L["Tier-L Reasoner
70B+/MoE"] CLS -->|"diagram / DICOM / UI screenshot"| V["Tier-V Multimodal"] CLS -->|"retrieval / rerank"| E["Tier-E Embed/Rerank"] S & M & L & V & E --> BUDGET{"Cost-budget check
(P6)"} BUDGET -->|"within budget"| SERVE["Serving plane"] BUDGET -->|"over budget"| DEGRADE["Downshift tier or queue"] Routing inputs: task type, IEC 62304 class, context length, required latency SLO, remaining task cost budget, and the active fine-tuned LoRA adapter. The classifier itself is a cheap Tier-S model; its decision is logged as evidence so routing is auditable and reproducible. 2.4 Model-Serving Plane # Component Tech Responsibility vLLM PagedAttention, continuous batching High-throughput text generation, Tier-S/M/L Triton + TensorRT-LLM Compiled engines, in-flight batching Latency-critical / quantized serving KServe InferenceService CRDs, multi-LoRA hot-swap Standardized serving surface, canary/shadow, autoscale 2.5 Data / Knowledge Plane # Component Tech Responsibility Code Knowledge Graph Self-hosted property graph (Neo4j / JanusGraph / NebulaGraph) Symbols, call graph, requirement→code→test traceability Vector store Self-hosted (Qdrant / Milvus / Weaviate) ANN retrieval over code & docs using Tier-E embeddings Regulatory corpus Versioned doc store + full-text (OpenSearch) IEC 62304 / ISO 13485 / 14971 / Part 11 reference text Model Registry & Lineage MLflow Model versions, fine-tune lineage, signed digests, stage 2.6 Platform / Infra Plane # Component Tech Responsibility Cluster + GPU Kubernetes + NVIDIA GPU Operator (MIG) Scheduling, GPU lifecycle, driver/DCGM Distributed compute Ray + Kueue Training, batch eval, fan-out inference jobs Mesh & identity Istio + SPIFFE/SPIRE mTLS, workload identity, zero-trust east-west Secrets HashiCorp Vault Dynamic short-lived secrets, transit, PKI Autoscale KEDA Queue-driven scaling, scale-to-zero for idle tiers Delivery Argo CD + Argo Workflows GitOps, pipeline orchestration 2.7 Governance / Audit Plane # Component Tech Responsibility WORM evidence store Object store with object-lock (immutable), append-only ledger Part 11 records, trajectory dumps, gate verdicts Telemetry OpenTelemetry collectors → metrics/traces/logs backends End-to-end spans, cost attribution FinOps ledger Cost attribution service Per-task GPU-seconds, tokens, cost-per-green-PR Supply chain Sigstore/cosign + SLSA provenance + SBOM Signed images, models, and artifacts 3. End-to-end request / trajectory flow # sequenceDiagram autonumber participant Dev as Developer (IDE/CLI) participant RT as Agent Runtime participant KG as Knowledge/RAG participant GW as Routing Gateway participant MS as Model Serving participant SB as Sandbox participant POL as Policy Server participant EV as Eval Service participant AU as Autonomy Authz participant AUD as WORM Audit Dev->>RT: Intent (task, repo, requirement ID) RT->>AU: Request autonomy level (IEC 62304 class) AU-->>RT: Allowed level (Class C → dual-human required) RT->>KG: Assemble context (graph + ANN + full-text) KG-->>RT: Ranked context bundle (+provenance) RT->>GW: Generation request (task + context) GW->>GW: Classify → select tier/LoRA → budget check (P6) GW->>MS: Route to tier MS-->>RT: Candidate artifact (diff/code/tests) RT->>SB: Hermetic build + test (Verify) SB-->>RT: Build/test results alt verification fails RT->>GW: Repair request (failure context) GW->>MS: Re-generate (Repair) MS-->>RT: Revised artifact RT->>SB: Re-verify end RT->>POL: Policy gate (writes, licenses, secrets) POL-->>RT: Allow / Deny + rationale RT->>EV: Eval gate (deterministic suite) EV-->>RT: Pass/Fail (≥99.9% threshold) RT->>Dev: HITL review (class-proportional) Dev-->>RT: Approve / reject (dual for Class C) RT->>Dev: Open PR (signed) RT->>AUD: Write immutable evidence (P4, Part 11) Each step emits a span and a content-addressed evidence record. The trajectory is fully reconstructable from the audit store, satisfying P4 and 21 CFR Part 11. 4. Kubernetes deployment topology # flowchart TB subgraph CTRL["Control / CPU node pool"] NS_AGENT["ns: agent-runtime"] NS_HARNESS["ns: harness-control"] NS_KNOW["ns: knowledge"] NS_GOV["ns: governance-audit"] NS_PLAT["ns: platform (Vault, Istio, Argo)"] end subgraph GPU["GPU node pools"] POOL_S["pool: reflex (MIG 1g.10gb / L4)"] POOL_M["pool: worker (A10/L40S)"] POOL_L["pool: reasoner (H100/H200, NVLink)"] POOL_TRAIN["pool: train/batch (Kueue-managed)"] end subgraph SANDBOX["Sandbox node pool (CPU, isolated)"] NS_SB["ns: sandbox-exec (gVisor/Kata, no egress)"] end NS_HARNESS -->|route| POOL_S & POOL_M & POOL_L NS_AGENT --> NS_SB POOL_TRAIN --- RAY["Ray + Kueue queues"] 4.1 Namespaces # Namespace Contents Network policy agent-runtime Agent pods, A2A bus Egress only to MCP, gateway, knowledge harness-control Gateway, policy server, autonomy authz, eval Egress to serving + knowledge; ingress from agents model-serving-{s,m,l,v,e} vLLM/Triton/KServe per tier Ingress only from gateway knowledge KG, vector store, OpenSearch, MLflow Ingress from agents/harness sandbox-exec gVisor/Kata pods Default-deny all egress ; ephemeral governance-audit WORM store, OTel, FinOps Append-only ingest platform Vault, Istio control plane, Argo, SPIRE Cluster-internal 4.2 Node pools, MIG, and scheduling # Pool Hardware (example) MIG Scaling reflex (Tier-S) L4 / A10 1g.10gb partitions for high pod density KEDA, scale-to-zero off-hours worker (Tier-M) L40S / A10 optional MIG KEDA queue-driven reasoner (Tier-L) H100/H200, NVLink + GPUDirect full GPU, tensor/pipeline parallel conservative; warm pool ≥1 train/batch H100 multi-node full GPU Kueue gang-scheduling, preemptible GPU Operator manages drivers, DCGM exporters, MIG geometry, and time-slicing where MIG is too coarse. Kueue provides quota-managed queues for training and batch eval, with ClusterQueue / LocalQueue and gang scheduling for multi-node Tier-L fine-tunes. KEDA scales serving deployments off MCP/gateway queue depth and supports scale-to-zero for idle Tier-V/Tier-L adapters — central to P6. NetworkPolicies enforce default-deny; sandbox namespace is fully air-gapped from cluster services and the internet. 4.3 Multi-cluster, sovereign-VPC, and air-gap # flowchart LR subgraph SOV["Sovereign region cluster"] direction TB PROD["prod (serving + audit)"] VAL["validation"] end subgraph DEV["Dev cluster"] DEVNS["dev / experimentation"] end MIRROR["Artifact mirror
(images · models · pkgs)"] DEV -. promote (signed) .-> VAL VAL -. promote (signed) .-> PROD MIRROR --> SOV MIRROR --> DEV Sovereign-VPC: prod and validation run in a customer-controlled region/VPC; no cross-border data flow. Multi-cluster: dev separated from validation/prod clusters; promotion is signed-artifact-only (cosign verified at admission). Air-gap option: every dependency (base images, model weights, OS packages, eval datasets) is mirrored into an internal registry. No runtime reaches the public internet. The architecture has no hard internet dependency at request time (C5/P7). 5. Model-serving subsystem in depth # 5.1 Tier → hardware mapping # Tier Models (open-weight, fine-tuned) Hardware Serving Quant Tier-S "Reflex" 1-8B Qwen2.5-Coder-1.5B/7B, Llama-3.2-3B L4 / MIG slice Triton+TRT-LLM FP8 / INT8 Tier-M "Worker" 14-34B Qwen2.5-Coder-32B, StarCoder2-15B, DeepSeek-Coder-V2-Lite L40S / A10 vLLM FP8 / AWQ-INT4 Tier-L "Reasoner" 70B+/MoE Llama-3.3-70B, Qwen2.5-72B, DeepSeek-V3/R1-distill, Mixtral H100/H200 NVLink vLLM / TRT-LLM FP8 / GPTQ Tier-V "Multimodal" Qwen2.5-VL, Llama-3.2-Vision, InternVL, Pixtral L40S/H100 vLLM FP8 Tier-E "Embed/Rerank" bge, gte, jina-code, nomic L4 / CPU Triton / TEI INT8 5.2 Serving techniques # Technique Applied where Purpose Continuous / in-flight batching vLLM, Triton Throughput; amortize GPU (P6) PagedAttention KV-cache vLLM Memory efficiency, longer context Speculative decoding Tier-L with Tier-S drafter Lower latency on reasoner Quantization (FP8/AWQ/GPTQ) all tiers Fit larger models, more density Multi-LoRA hot-swap KServe/vLLM Many fine-tuned adapters per base; per-task adapter without reload Tensor/pipeline parallel Tier-L Serve 70B+/MoE across GPUs 5.3 Lifecycle: autoscale, hot-swap, canary/shadow # flowchart LR REGY["MLflow registry
(signed digest)"] -->|promote| KS["KServe InferenceService"] KS --> CANARY["Canary 5%"] KS --> STABLE["Stable 95%"] SHADOW["Shadow (mirror, no user impact)"] -.eval.-> EVAL["Eval Service"] EVAL -->|pass ≥99.9%| ROLL["Promote canary→stable"] EVAL -->|fail| HALT["Halt + rollback"] New adapters/models enter as shadow (traffic mirrored, outputs eval'd offline), then canary (small %), then stable — each gated by the deterministic eval suite. KEDA autoscales each tier on queue depth; idle adapters scale to zero, base engines retain a warm minimum. Every promotion verifies a cosign signature and a pinned model digest so prod inference is reproducible (P2, P7). 6. Data & knowledge subsystem # flowchart TB SRC["Sources: repos · requirements · DHF · regs"] --> ING["Ingestion + sanitization
(PII/secret scrub, license tag)"] ING --> KG["Code Knowledge Graph"] ING --> EMB["Embeddings (Tier-E)"] ING --> FTS["Full-text index (OpenSearch)"] EMB --> VEC["Vector store"] QRY["Retrieval orchestrator"] --> KG QRY --> VEC QRY --> FTS KG & VEC & FTS --> FUSE["Fusion + rerank (Tier-E)"] FUSE --> CTX["Context bundle + provenance"] 6.1 Ingestion & sanitization # Sources (repos, requirements/DHF, regulatory corpus) pass through ingestion that scrubs secrets/PII , tags license and IEC 62304 class, and content-addresses each chunk so retrieval is reproducible and auditable. 6.2 Retrieval modes # Mode Backend Use Graph traversal Code Knowledge Graph (self-hosted property graph) Call graph, requirement→code→test traceability, blast-radius ANN (semantic) Vector store (Qdrant/Milvus) Similar code, prior solutions, doc semantics Full-text (lexical) OpenSearch Exact symbols, error strings, reg clauses Results are fused and reranked by a Tier-E model. The provenance of every retrieved chunk (source, version, digest) travels with the context bundle into the trajectory and the audit record (P4). The code knowledge graph is the self-hosted analogue of a managed graph-of-code service (Spanner-graph-class); it must be operable inside the sovereign/air-gapped boundary. 7. Control / governance plane # 7.1 Evidence production (P4 / Part 11) # Every plane emits structured evidence to the governance plane: Evidence Producer Stored Intent + autonomy decision Agent runtime + autonomy authz WORM Context bundle + provenance Knowledge plane WORM Routing/classification decision Gateway WORM Model digest + LoRA used + tokens/GPU-s Serving + FinOps WORM + ledger Sandbox build/test results Sandbox WORM Policy verdict + rationale Policy server WORM Eval verdict + dataset digest Eval service WORM HITL approver identity + signature Review UI WORM Records are written to an immutable, append-only, object-locked (WORM) store, time-stamped and attributable, satisfying 21 CFR Part 11 electronic-records/signatures and GAMP 5 traceability. 7.2 Runtime policy & autonomy enforcement (P3) # flowchart LR ACT["Agent action / tool call"] --> PEP["MCP enforcement point"] PEP --> OPA["Policy server (Rego)"] PEP --> AZ["Autonomy authz (IEC 62304 class)"] OPA -->|deny| BLOCK["Block + evidence"] AZ -->|Class C| DUAL["Require dual human control"] AZ -->|Class A/B| LEVEL["Apply allowed autonomy"] OPA -->|allow| EXEC["Execute"] DUAL --> EXEC LEVEL --> EXEC Policy is deterministic (Rego bundles, versioned, signed) wrapping the probabilistic agent (P2). Autonomy is risk-proportional : Class A/B may allow higher automation; Class C always requires dual human control before any write/promote. Identity for every actor (human or workload) is SPIFFE/SPIRE-issued; secrets are short-lived Vault leases. 8. Reference environments & promotion # Environment Purpose Models Data Gate to next dev Experimentation, adapter dev latest candidate LoRAs synthetic/masked unit + shadow eval pass validation Formal V&V (CSA/GAMP 5) release-candidate, pinned digests masked production-like full deterministic suite ≥99.9%, signed prod Live agentic SDLC only validated, signed models real (sovereign) n/a flowchart LR DEV["dev"] -->|signed artifact + eval pass| VAL["validation"] VAL -->|full V&V + cosign verify| PROD["prod"] PROD -. rollback (pinned prior digest) .-> PROD Promotion is GitOps + signed-artifact only : Argo CD reconciles only cosign-verified images and MLflow-registered model digests; admission control (Gatekeeper) rejects anything unsigned or off-registry. In air-gapped/regulated networks, promotion crosses the boundary as a signed, mirrored bundle — never a live pull. 9. Build-vs-buy and self-hosting rationale # Concern Decision Rationale (ties to 08 ) LLM inference Build/host (open-weight) C1/P7: sovereignty, reproducibility, no PHI/IP egress; predictable cost-per-green-PR Model fine-tuning Build (Ray+Kueue) Domain/device-specific quality; full lineage in MLflow Orchestration & harness Build P5: the harness is the differentiator and the 99.9% lever Serving runtime Buy/adopt OSS (vLLM/Triton/KServe) Mature, self-hostable, no SaaS lock-in Knowledge graph / vector / search Buy/adopt OSS, self-host Operable in air-gap; avoids managed-SaaS data residency issues Identity/secrets/mesh Adopt OSS (SPIRE/Vault/Istio) Zero-trust standard, self-hostable Supply chain Adopt OSS (Sigstore/SLSA) Required for §524B / reproducibility The economic case (GPU amortization via batching, MIG density, scale-to-zero, tiered routing) is developed in 08-token-and-gpu-economics.md . The architecture is intentionally biased toward owning the harness and hosting the models , and adopting mature OSS for undifferentiated platform layers. 10. Architecture-to-maturity mapping (ASMM-Med) # Which components must be operational at each level (see 02-maturity-model.md ): Component / capability L1 Governed Assistance L2 Spec-Driven Bounded L3 Orchestrated Agentic L4 Validated Autonomous L5 Self-Optimizing Self-hosted serving (vLLM/Triton) ● ● ● ● ● Tiered fleet + routing gateway ◐ ● ● ● ● Multi-LoRA hot-swap — ◐ ● ● ● Knowledge plane (KG+vector+FTS) ◐ ● ● ● ● Sandbox verify (Generate→Verify) — ● ● ● ● Deterministic eval gate (≥99.9%) — ◐ ● ● ● Policy server + autonomy authz (P3) ◐ ● ● ● ● Agent runtime + MCP tool plane — ◐ ● ● ● A2A multi-agent — — ● ● ● WORM evidence / Part 11 (P4) ◐ ● ● ● ● FinOps cost-per-green-PR (P6) — ◐ ● ● ● Canary/shadow + auto-promotion — — ◐ ● ● Closed-loop self-optimization — — — ◐ ● Legend: ● required · ◐ partial/emerging · — not yet. Adoption sequencing for these capabilities is detailed in 09-adoption-roadmap.md . Appendix A — Component-to-principle traceability # Component P1 P2 P3 P4 P5 P6 P7 Routing gateway ● ● ● ● ● ● Sandbox verify ● ● ● ● ● Eval gate ● ● ● ● ● Policy + autonomy authz ● ● ● ● ● WORM evidence store ● ● ● Model serving (self-hosted) ● ● ● ● ● ● FinOps ledger ● ● End of document — 03-reference-architecture.md ← Previous 02 · Maturity Model Next → 04 · Model Strategy & Fine-Tuning On this page 1. Architecture goals, constraints, and the layered view 2. Logical components per plane 3. End-to-end request / trajectory flow 4. Kubernetes deployment topology 5. Model-serving subsystem in depth 6. Data & knowledge subsystem 7. Control / governance plane 8. Reference environments & promotion 9. Build-vs-buy and self-hosting rationale 10. Architecture-to-maturity mapping (ASMM-Med)
# 04 · Model Strategy & Fine-Tuning · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/04-model-strategy-and-finetuning.html
## 04 — Model Strategy & Fine-Tuning #
04 · Model Strategy & Fine-Tuning · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate 04 — Model Strategy & Fine-Tuning # Figure E — Tiered Model-Fleet Routing · open SVG Document set: Agentic-Native SDLC for Regulated Medical Device Engineering Status: Controlled engineering reference · Revision date: May 2026 Owning function: ML Platform & Quality Engineering Cross-references: 01-requirements · 02-maturity-model · 03-reference-architecture · 05-evaluation-and-validation · 06-agentic-workflows · 07-security-and-compliance · 08-token-and-gpu-economics · 09-adoption-roadmap Regulatory anchors: IEC 62304 · ISO 13485 / QMSR · ISO 14971 · FDA CSA · GAMP 5 · 21 CFR Part 11 · ISO/IEC 42001 · FDA AI-enabled device guidance + PCCP Note on thresholds: All numeric thresholds in this document (e.g., ≥99.9%, abstention rates, eval gates) are placeholders pending calibration in 05 . They denote intent and governance structure, not yet-ratified acceptance criteria. 1. Strategy rationale: why a tiered, multi-model fleet of open-weight models # The toolchain is the regulated artifact, not the device. Per Principle P5 (the harness is the product) , the models are components inside a verification harness whose system-level correctness target is ≥99.9% at the release gate ( P1 ). No single model — however large — is the source of that guarantee; the guarantee emerges from Generate → Verify → Repair → Gate loops where deterministic checks ( P2: determinism wraps probabilism ) bound probabilistic generation. Given that, model strategy optimizes for fitness-per-task at lowest defensible cost , not for a single frontier model. 1.1 Why tiered multi-model instead of one big model # Driver One big model Tiered fleet (chosen) Cost (P6: cost-per-green-PR) Every autocomplete keystroke pays 70B+ inference. Economically fatal at 1000+ devs. Reflex-tier 1–8B handles the high-frequency 80%; Reasoner-tier reserved for rare hard plans. Tie to 08 . Latency 70B autocomplete = unusable IDE latency. Tier-S sub-100ms-class on small GPUs; interactive paths never touch large models. Specialization Generalist regresses on niche tasks (embedded C, DICOM, regulatory drafting). Per-domain LoRA adapters (§5) sharpen narrow tasks without retraining a monolith. Abstention & calibration (§9) Monolith over-commits; one calibration curve for all tasks. Per-tier/per-task calibration; cheap router/classifier abstains early and escalates. Blast radius / governed evolution (§8) One promotion changes everything; PCCP scope is the whole org. Adapter-scoped change control; canary one adapter without re-validating the fleet. GPU packing Coarse, wasteful allocation. Multi-LoRA hot-swap on shared base weights; bin-pack tiers across the cluster. Determinism (P2) Harder to wrap one opaque large model in deterministic checks. Small, fast verifiers and routers are themselves deterministic-friendly. The fleet is a portfolio : route each task to the smallest capable model (§10), escalate on low confidence, and let deterministic verifiers — not model scale — own the correctness guarantee. 1.2 Why open-weight + self-hosted (P7: self-hosted, sovereign, reproducible) # Requirement Why SaaS LLM APIs are disqualified What open-weight self-host gives us Reproducibility (P7, critical) Vendor silently updates the model; yesterday's evidence is not reproducible. We pin exact weights + tokenizer + config by digest; a fine-tune is reproducible byte-for-byte. CSA / Part 11 evidence (P4) No control over model lineage; cannot sign the chain. Full signed lineage dataset→base→adapter→eval→registry (§7) becomes defensible audit evidence. Data sovereignty (PHI/PII/IP) Source code, design history, and PHI-adjacent data leave the boundary. All training and inference stay inside the K8s/VPC trust boundary (see 07 ). Determinism & control No control over sampling, decode, or version. We fix seeds, decode params, and serving stack (vLLM/Triton). Cost (P6) Per-token vendor pricing scales with org size; opaque. Amortized GPU cost, first-class and measurable ( 08 ). Longevity Models deprecated by vendor on vendor timeline. We retain weights indefinitely for re-validation and regulatory defense. HARD CONSTRAINT (restated): Self-hosted, fine-tuned, open-weight models only. No SaaS LLM APIs anywhere in the SDLC toolchain. 2. The fleet specification # Five tiers. All tiers are served from the shared stack: K8s + NVIDIA GPU Operator , vLLM / Triton + TensorRT-LLM / KServe with multi-LoRA hot-swap . See 03 for serving topology. Tier Name Params Example base models (open-weight) Context window Serving HW (per replica) Quantization Typical tasks Why this tier Tier-S Reflex 1–8B Qwen2.5-Coder-1.5B / 7B, Llama-3.2-3B 8–32k 1× L4 / A10 (or MIG slice) FP8 / INT8; AWQ/GPTQ-4bit for 1.5B Autocomplete, classify, route, redact/PII-scrub, abstain-or-escalate gate High-frequency, low-latency, cost-dominant path. Must be cheap (P6) and fast. Tier-M Worker 14–34B Qwen2.5-Coder-32B, StarCoder2-15B, DeepSeek-Coder-V2-Lite 16–128k 1–2× A100/H100 80GB FP8; AWQ-4bit option Test generation, refactor, code review, doc drafting, structured edits Workhorse for bounded automation (ASMM-Med L2/L3 ). Strong code with feasible cost. Tier-L Reasoner 70B+ / MoE Llama-3.3-70B, Qwen2.5-72B, DeepSeek-V3 / R1-distill, Mixtral-8x22B 32–128k 2–8× H100 80GB (TP/PP) FP8; INT4 for offline Architecture, multi-step planning, hard root-cause, spec decomposition Rare, high-value reasoning. Reserved; never on interactive hot paths. Tier-V Multimodal 7–90B Qwen2.5-VL, Llama-3.2-Vision, InternVL2, Pixtral 8–128k 1–4× A100/H100 80GB FP8 / AWQ Parse design PDFs, schematics, UI screenshots, imaging artifacts, diagram→spec Inputs in this domain are visual (design history files, DICOM-adjacent). Tier-E Embedding / Rerank 0.1–1.5B bge, gte, jina-code, nomic-embed 512–8k 1× L4 / A10 (CPU fallback) FP16 / INT8 RAG retrieval, code search, dedup, contamination detection, rerank Retrieval substrate for every agentic workflow ( 06 ). Routing summary. A Tier-S classifier/router triages every request: trivial → answer; ambiguous/high-risk → escalate to Tier-M/L; visual → Tier-V; retrieval → Tier-E first. Routing policy is governed by IEC 62304 class of the affected artifact ( P3: autonomy by class A/B/C ) and recorded as evidence ( P4 ). 3. Base-model selection criteria & governance # A base model is a supplier-provided component under ISO 13485 / QMSR supplier controls and GAMP 5 categorization. No base model enters the fleet without passing the gate below and landing in the Approved Base-Model Registry . 3.1 Selection criteria # Criterion Requirement Evidence captured License compatibility License must permit commercial + regulated use, self-hosting, fine-tuning, and redistribution of derivatives internally. Legal sign-off mandatory (see §3.2). License text, SPDX id, legal approval record. Provenance Weights obtained from the authoritative publisher; digest verified. No re-uploads of unknown origin. Source URL, publisher identity, SHA-256 of weights + tokenizer. Security scan of weights Scan serialized weights for unsafe deserialization (reject pickle where possible; require safetensors ), embedded code, and known-bad artifacts. Quarantine until clean. Scan report, scanner version, verdict. Model card Documented training data summary, intended use, known limitations, eval baselines, and bias notes. Missing card → not approved. Stored model card + internal addendum. Capability baseline Passes minimum task-suite scores in 05 before any fine-tuning. Eval run id, scores. Maintainability Supported by serving stack (vLLM/TensorRT-LLM), tokenizer stable, reasonable VRAM footprint. Compatibility matrix entry. 3.2 License review (open-weight ≠ unrestricted) # "Open-weight" describes weight availability, not unrestricted rights. Each license is reviewed individually; the table below is engineering guidance, not legal advice — Legal sign-off per model is mandatory and recorded in the registry. License family Typical examples Commercial / regulated use Watch-outs (review per version) Apache-2.0 / MIT Qwen2.5 (most sizes), StarCoder2, many bge/gte Generally permissive Confirm the specific checkpoint's license; some variants differ. Llama Community License Llama-3.x family Permitted with conditions Acceptable-use policy, attribution/naming requirements, large-MAU clause. Model-specific bespoke DeepSeek, some VL models Case-by-case Field-of-use, output/derivative terms, redistribution limits. Non-commercial / research-only Some checkpoints Disqualified Never admitted to the production fleet. 3.3 Approved Base-Model Registry # Maintained in the MLflow registry with signed entries. Schema: Field Example model_uid base/qwen2.5-coder-32b weights_digest sha256:… (safetensors) tokenizer_digest sha256:… license_spdx / legal_approval_ref Apache-2.0 / LGL-2026-0142 provenance_url / publisher authoritative source security_scan_ref / verdict SCAN-2026-0331 / clean model_card_ref stored card + addendum tier / serving_compat Tier-M / vLLM,TRT-LLM approval_state approved \ cosign_signature Sigstore/cosign over the manifest Only approved bases may be parents of a fine-tune. Revocation propagates to all derived adapters (§8). 4. The fine-tuning pipeline # flowchart TD subgraph SRC["Sourced & governed inputs (§6)"] A1[Internal sanitized corpus
code · docs · tickets] A2[Task / instruction datasets] A3[Preference pairs
chosen / rejected] A4[Verifier-filtered synthetic data] end B[("Approved Base-Model
Registry (§3)")] --> S1 A1 --> S1["Stage 1: DAPT
Domain-Adaptive Continued Pretraining
(usually full FT or large LoRA)"] S1 --> S2["Stage 2: SFT
Instruction / task tuning
(LoRA or full FT)"] A2 --> S2 A4 --> S2 S2 --> S3["Stage 3: Preference alignment
DPO / ORPO (TRL)"] A3 --> S3 S3 --> S4["Stage 4: Task/Domain LoRA adapters
(PEFT/QLoRA) — one per specialization (§5)"] S4 --> E["Eval gate (05)
deterministic suites + abstention checks"] E -->|pass| R[("MLflow registry
signed adapter + lineage (§7)")] E -->|fail| X[Repair / re-tune / reject] R --> SERVE["Multi-LoRA serving
vLLM/Triton hot-swap on shared base"] classDef gate fill:#eef,stroke:#446; class E gate; 4.1 When to use each stage # Stage Purpose Method Use when Skip when 1 — DAPT (Domain-Adaptive Continued Pretraining) Inject domain distribution (embedded C idioms, regulatory register, internal APIs) Continued pretraining on the sanitized internal corpus ; full FT or large-rank LoRA; DeepSpeed/FSDP via Ray+Kueue Base is unfamiliar with the domain vocabulary/style at the token level Base already strong in-domain; only behavior shaping needed 2 — SFT (Supervised Fine-Tuning) Teach task format & instruction following (test-gen schema, review rubric, doc templates) TRL SFT; LoRA/QLoRA usually sufficient; full FT only if LoRA underfits Almost always — this is the primary lever for task behavior Task is purely retrieval/format-trivial 3 — Preference alignment Shape preferences : prefer abstention over guessing, prefer compiling code, prefer cited regulatory claims DPO / ORPO (TRL) on chosen/rejected pairs Need to suppress over-confidence, hallucination, or unsafe patterns (§9) No reliable preference signal yet 4 — Task/Domain LoRA Narrow, swappable specialization PEFT LoRA/QLoRA adapters on the aligned base Per-domain (§5) capability needed without forking the base One general adapter already meets the eval gate 4.2 LoRA vs full fine-tuning — decision rule # Use LoRA / QLoRA when… Use full FT when… Behavior/format adaptation on a capable base (most SFT, all per-domain adapters) DAPT requires moving the base distribution substantially You need many swappable specializations on shared weights (multi-LoRA serving) Tokenizer/vocab must change GPU/cost budget is tight (P6); QLoRA fits on fewer GPUs LoRA repeatedly underfits the eval target after rank/data tuning Fast iteration and small, signable artifacts are required A new long-lived base derivative is justified and will itself enter the registry Default posture: prefer LoRA. Full FT is the exception and requires a documented justification plus its own registry entry as a derived base. 5. Domain specialization # Each domain ships as a named LoRA adapter over an approved (optionally DAPT'd) base, independently versioned, eval-gated, and signed. Multi-LoRA serving hot-swaps the right adapter per request. Domain adapter Tier(s) Specialized capability Example tasks embedded-fw-c M / L Embedded/firmware C, MISRA-style constraints, ISRs, fixed-point, no-malloc patterns Generate/refactor firmware, flag undefined behavior, MISRA review imaging-pipeline M / V Imaging processing pipelines, array/tensor ops, numerical stability Pipeline code-gen, perf refactor, artifact reasoning dicom-adjacent M / V DICOM-adjacent metadata, header semantics, de-ID conventions Parse/validate metadata, generate handling code reg-doc-drafting M / L Regulatory register; IEC 62304 / ISO 14971 / Part 11 phrasing; traceable claims Draft design history items, risk entries, SOUP rationale test-generation M Coverage-oriented unit/integration test synthesis with assertions Generate tests to push coverage and mutation score code-review M Project-specific review rubric, severity classification Structured review with cited rule ids router-classify S Triage, risk/class tagging, escalation, redaction Route + abstain-or-escalate gate 5.1 Multimodal angle (Tier-V) # Design inputs in this domain are inherently visual. Tier-V adapters target: Design PDFs / design history files → extract structured requirements/specs (feeds 06 ). Schematics / block diagrams → derive interfaces, signal lists, architecture facts. UI screenshots → verify UI against spec; detect drift. Imaging artifacts → describe/triage visual anomalies (advisory only; never a clinical claim). All Tier-V outputs are advisory inputs to deterministic verifiers , not autonomous decisions; class-C-affecting outputs always require human confirmation ( P3 ). 6. Data strategy for fine-tuning # Training data is a controlled, versioned, signed artifact . The dataset is as much a regulated input as the model. Concern Control Sourcing Internal code, design docs, tickets/issues, review history — pulled via governed connectors with access controls ( 07 ). PII / PHI scrubbing Mandatory de-identification before any data leaves the source boundary into training. Multi-pass: pattern + Tier-S redaction model + human spot-audit. No PHI in training sets, ever. IP / license hygiene Exclude third-party code with incompatible licenses; track provenance per record; quarantine unknown-origin snippets. Dataset versioning & signing Immutable, content-addressed dataset snapshots ( dataset_uid + digest), registered in MLflow, cosign-signed . Train/test contamination control Use Tier-E embeddings to detect near-duplicates between training data and held-out eval sets ( 05 ); fail the build on overlap above threshold. Eval sets are sealed and never enter training. Synthetic data Generated by Tier-M/L, then verifier-filtered : only synthetic examples whose outputs pass deterministic checks (compiles, tests pass, schema-valid) are retained. Unverifiable synthetic data is discarded. Provenance labels Every record tagged internal / synthetic-verified / public-permissive for auditability and ablation. Contamination is a correctness-and-evidence risk, not a metric nuisance. A model trained on its own eval set produces inflated scores that cannot support a defensible ≥99.9% claim ( P1/P4 ). Contamination control is a release gate (§11 anti-pattern). 7. Reproducibility & validation as first-class (P7) # A fine-tune must be reproducible and defensible as CSA evidence . We lock every input and sign the full lineage so an auditor can re-derive the artifact. 7.1 What is locked # Locked input Mechanism Dataset Content-addressed snapshot ( dataset_uid + digest), signed (§6). Base model weights_digest + tokenizer_digest from the Approved Registry (§3). Config Hyperparameters, stage sequence, LoRA rank/targets, decode params — versioned YAML (Axolotl/Llama-Factory/torchtune), digested. Seeds All RNG seeds (data shuffling, init, dropout) pinned. Environment Container image digest, CUDA/driver, library versions (PEFT/TRL/DeepSpeed) recorded. Eval Eval suite version + sealed test-set digest ( 05 ). 7.2 Signed lineage chain # flowchart LR D["dataset_uid
(signed digest)"] --> A BM["base model_uid
(registry digest)"] --> A CFG["config + seeds + env
(digest)"] --> A A["adapter / model artifact
(SHA-256)"] --> EV EV["eval run
(suite ver + scores)"] --> REG REG[("MLflow registry entry
cosign-signed, SLSA provenance")] Each edge is a cosign attestation ; the registry entry carries SLSA provenance . The chain answers the auditor's question — "show me exactly how this model was produced and that nothing changed" — and is the artifact-level realization of Part 11 (P4) and CSA . 7.3 How this satisfies P7 and D2-L4 # P7 (reproducible): Any registered fine-tune is byte-reproducible from locked inputs; re-running the pipeline yields the same artifact digest (modulo documented nondeterminism, which is itself bounded and recorded). D2-L4 ( 02 ): Validated Autonomous Agents require validated models . The signed lineage + eval-gated promotion (§8) is the model-side evidence package that lets an agent operate autonomously within its IEC 62304 class envelope. 8. Model lifecycle & governed evolution # Models evolve under a PCCP-style predetermined change control applied to the toolchain models (reusing the FDA PCCP concept; the device itself is governed separately). Promotion is eval-gated ( 05 ); no promotion bypasses the gate. flowchart LR C["candidate
(new adapter / base)"] --> SH["shadow
(mirror traffic, no effect)"] SH --> CN["canary
(small % real, guarded)"] CN --> PR["promote
(default for tier/domain)"] PR --> DP["deprecate
(retain weights + lineage)"] SH -->|fail gate| RJ[reject] CN -->|regression| RB[rollback] Phase Gate / exit criteria Evidence Candidate Lineage signed (§7); passes offline eval suite + abstention/calibration checks (§9) Registry entry, eval run id Shadow Mirrored traffic; no regression vs incumbent on live distribution; no safety violations Shadow comparison report Canary Bounded % of real traffic by IEC 62304 class (lower class first); cost-per-green-PR within budget (P6) Canary metrics, guardrail logs Promote Meets/exceeds incumbent on all gated metrics; sign-off recorded Promotion record, signatures Deprecate Successor promoted; weights and lineage retained for re-validation/defense Retention record 8.1 PCCP-style predetermined change control (toolchain models) # The Predetermined Change Control Plan for the model fleet specifies, in advance : the allowed change types (e.g., new domain adapter, refreshed SFT data), the fixed eval protocol that gates them, the rollback triggers, and the autonomy class affected. Changes inside the envelope flow through the lifecycle without re-opening the whole validation; changes outside it require plan revision. This ties to D1-L5 ( 02 ) (self-optimizing under governance) and the evaluation regime in 05 . Base-model revocation (§3) forces immediate deprecation of all derived adapters. 9. Abstention & calibration # The ≥99.9% system property ( P1 ) depends on models that decline rather than guess on out-of-distribution or high-risk inputs, escalating to a larger tier or a human. Abstention is a trained and served behavior, validated in 05 . Mechanism Where Effect Preference training for abstention Stage 3 DPO/ORPO (§4) Prefer "insufficient evidence → escalate" over a confident wrong answer Calibrated confidence Tier-S router + per-task heads Confidence thresholds tuned so high-confidence ≈ high-accuracy Abstain-or-escalate gate Serving (router) Below threshold → escalate tier or hand to human; never silently proceed Selective prediction metrics Eval ( 05 ) Track coverage vs risk; gate on risk at fixed coverage , not raw accuracy Class-aware strictness (P3) Routing policy IEC 62304 class C → conservative thresholds, mandatory human confirmation Calibration target: in the operating region, a high-confidence answer is correct ≥99.9% ; everything else abstains and routes to verification or a human. A miscalibrated-but-accurate model is not acceptable — abstention behavior is itself an eval gate. 10. Right-sizing & cost linkage # Smallest-capable-model principle (P6): route every task to the smallest model that passes the gate; escalate only on abstention. Lever Action Cost effect (→ 08 ) Tiered routing Tier-S handles the high-frequency majority; Tier-L is rare Largest single cost lever; collapses per-token spend Distillation Capture Tier-L behavior (traces, preferences) → train Tier-M/S adapters Moves capability down a tier at a fraction of inference cost Quantization FP8/AWQ/GPTQ per tier (§2) More replicas per GPU; lower latency Multi-LoRA hot-swap Many domain adapters on one shared base Eliminates per-domain base replicas; high GPU packing Abstention budgeting Escalate only when justified (§9) Prevents needless large-model calls Right-sized context Use the minimum context window that passes eval Lower KV-cache cost Distillation note. The verifier-filtered synthetic pipeline (§6) is the distillation substrate: only Tier-L outputs that pass deterministic verification become training data for smaller adapters, so distillation transfers verified behavior, not hallucinations. The primary metric for right-sizing decisions is cost-per-green-PR (P6) , owned in 08 . 11. Anti-patterns # # Anti-pattern Why it fails here Required control A1 Unversioned models No reproducibility, no defensible evidence; violates P7/P4 Every model is a signed registry entry with full lineage (§7) A2 Training on the eval set Inflated scores cannot support ≥99.9% ( P1 ); fraudulent evidence Embedding-based contamination gate; sealed eval sets (§6, 05 ) A3 Unscanned weights Unsafe deserialization / supply-chain compromise Mandatory weight scan + safetensors before approval (§3) A4 License violation Legal and regulatory exposure; non-commercial weights in production Per-model legal sign-off in the registry (§3.2) A5 SaaS LLM API "just for this one thing" Breaks sovereignty, reproducibility, Part 11 chain ( P7 ) Hard constraint: open-weight self-host only (§1.2) A6 One big model for everything Cost-fatal, latency-fatal, coarse governance (§1.1) Tiered fleet + smallest-capable routing (§2, §10) A7 Over-confident models (no abstention) Confident wrong answers break the 99.9% property Abstention training + calibration gates (§9) A8 Promotion without eval gate Unvalidated change reaches users; violates D2-L4 Eval-gated candidate→shadow→canary→promote (§8) A9 PHI/IP in training data Privacy/IP breach; non-compliant corpus Mandatory scrubbing + provenance labels (§6, 07 ) A10 Unverified synthetic data Trains models on hallucinations; degrades correctness Verifier-filtering only; discard unverifiable (§6) Summary # The model strategy is a portfolio of small, specialized, signed, open-weight models governed as regulated components: tiered for cost/latency/specialization, fine-tuned through a locked-and-signed pipeline (DAPT → SFT → DPO/ORPO → LoRA), specialized per domain via hot-swappable adapters, and promoted only through eval-gated, PCCP-style change control. The correctness guarantee lives in the harness (P5) and its deterministic verification ( P2 ), not in model scale; the models contribute calibrated capability and disciplined abstention (P1, §9) , and reproducible, signed lineage (P7, §7) is what makes every fine-tune defensible CSA evidence. Implementation detail for evaluation lives in 05 and for economics in 08 . ← Previous 03 · Reference Architecture Next → 05 · Evaluation & Validation On this page 1. Strategy rationale: why a tiered, multi-model fleet of open-weight models 2. The fleet specification 3. Base-model selection criteria & governance 4. The fine-tuning pipeline 5. Domain specialization 6. Data strategy for fine-tuning 7. Reproducibility & validation as first-class (P7) 8. Model lifecycle & governed evolution 9. Abstention & calibration 10. Right-sizing & cost linkage 11. Anti-patterns
# 05 · Evaluation & Validation · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/05-evaluation-and-validation.html
## 05 — Evaluation & Validation: Earning 99.9% #
05 · Evaluation & Validation · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate 05 — Evaluation & Validation: Earning 99.9% # Figure B — The 99.9% Release-Gate Assurance Pipeline · open SVG Part of Agentic-Native SDLC for Regulated Medical Device Engineering . Status: Reference baseline — May 2026. Thresholds shown are org-set placeholders pending QMS ratification. Audience: engineering, quality/QA, regulatory affairs, validation (CSV/CSA) owners. Cross-references: 01-requirements.md , 02-maturity-model.md , 03-reference-architecture.md , 04-model-strategy-and-finetuning.md , 06-agentic-workflows.md , 07-security-and-compliance.md , 08-token-and-gpu-economics.md , 09-adoption-roadmap.md . This document is the core of the 99.9% argument . Everything else in the program — the model fleet, the harness, the maturity ladder — exists to feed and be governed by the evaluation and validation system described here. The thesis is Principle P1 : 99.9% is a system property, not a model property. No open-weight model, fine-tuned or otherwise, will be claimed to be 99.9% correct. Instead we engineer a pipeline — Generate → Verify → Repair → Gate — whose aggregate release-gate correctness reaches the target, and we prove it with deterministic, recorded, reproducible evidence (P2, P4, P7). 1. Reframing the 99.9%: what number are we actually defending? # The single most common error in agentic-SDLC programs is conflating three different quantities under one "99.9%" banner. They are measured differently, owned differently, and validated differently. # Quantity Definition Who is responsible Do we engineer it? (a) Model accuracy Probability a single model invocation produces a correct artifact on a representative task distribution Model strategy (04) No — informative, never gated on directly (b) System release-gate correctness Probability that an artifact admitted past the release gate is correct, given all verifiers, repair, and HITL Harness + Eval (this doc) Yes — primary target ≥ 99.9% (c) Escape-rate Probability that a defect reaches a regulated artifact (released code, DHF/DHR record, submission content) undetected Whole SDLC + post-market Yes — minimized; the safety-relevant number State plainly: we do not promise model accuracy (a). We engineer system correctness (b) and we drive down escape-rate (c). A model that is only 92% accurate but abstains or fails loudly on the other 8% can feed a 99.9% system, because the system never admits the bad 8% — it routes it to repair, escalation, or human control. Conversely, a 99% model with weak gating can have a worse escape-rate than a 90% model with strong gating. The model is a generator ; the gate is the guarantee . 1.1 Metric definitions (precise) # Let the unit of evaluation be a change-set (a PR-sized artifact) or an eval item (a single golden task). Define a "correct" outcome against the deterministic ground truth for that item. Metric Formula Meaning / use Precision (of the gate) TP / (TP + FP) Of artifacts the gate passes , fraction actually correct. This is (b). Drives release-gate correctness. Recall (of the gate) TP / (TP + FN) Of correct artifacts, fraction the gate passes. A proxy for first-pass throughput; low recall = wasted generation, not a safety issue. Escape-rate (defects reaching a regulated artifact) / (total artifacts released) (c). The post-gate failure probability. Target is an org-set ceiling, e.g. ≤ 1e-3, tracked per safety class. Abstention-rate (abstained or IDK items) / (total items) How often the system declines and escalates rather than guessing. A yield lever , not a defect (§6.4). First-Pass Yield (FPY) (items passing all gates with zero repair iterations) / (total items) Generation quality + harness efficiency. An economic metric (P6, ties to 08), not a safety gate. Repair-adjusted yield (items passing after ≤ K repair iterations) / (total items) Throughput including the bounded repair loop. False-pass rate (FPR_gate) FP / (TP + FP + TN + FN) admitted The complement view of escape-rate at the gate boundary; the number a CSA risk assessment cares about most. A critical conceptual rule, established here and enforced throughout: precision and escape-rate are the safety metrics; recall and FPY are the economic metrics. When the two trade off, safety wins — we prefer a system that abstains and escalates (lower FPY, higher cost) over one that admits a marginal artifact (higher FPY, higher escape-rate). This is a direct expression of risk-proportional autonomy (P3). 2. The assurance pipeline # Every artifact produced by an agent — code, test, spec, config, document fragment — flows through the same assurance pipeline before it can become evidence or be released. Determinism dominates the critical path (P2): probabilistic checks may inform and escalate , but they may never be the sole admit/reject decision for Class B or Class C work. flowchart TD A[Intent / Spec
requirement, Gherkin, ticket] --> B[Generate
model fleet S/M/L/V/E
constrained + spec-conditioned decoding] B --> C{Deterministic
Verify Layer} C -- fail --> D[Repair Loop
bounded budget K
feed verifier output back] D --> B C -- pass --> E{Eval Gate
golden suites + thresholds
statistical acceptance} E -- below threshold --> D E -- secondary signal only --> F[LLM-as-judge / rubric
NON-gating for B/C
escalation trigger only] F --> G E -- pass --> G{HITL
risk-proportional
Class C = dual control} G -- changes requested --> D G -- abstain / IDK --> H[Escalate to human owner] G -- approved --> I[Release Record
signed evidence bundle
traceability + Part 11 audit] D -- budget exhausted --> H I --> J[(Production / DHF / DHR
regulated artifact)] J -. production failures captured .-> K[Regression Capture
new golden items -> §8] K -. closed loop L5 .-> E Stage Role Why it sits here Generate Produce candidate artifact(s); may emit N samples for self-consistency Probabilistic by nature; cheapest to make wrong loudly via structure constraints Deterministic Verify Hard, reproducible accept/reject from compilers, tests, analyzers, schema/policy checks The workhorse. Same input → same verdict, always (P7). This is what makes the gate auditable. Repair loop Feed verifier diagnostics back to the generator; retry within a bounded budget K Converts a probabilistic generator into a system that converges ; budget prevents infinite/cost-runaway loops Eval Gate Compare against versioned golden suites + statistical acceptance criteria by safety class Decides admit/reject on evidence , not vibes; applies confidence-interval thresholds (§6) HITL Human review proportional to IEC 62304 class; Class C always dual control (P3) Eval score never substitutes for required human judgment on high-risk artifacts Release record Emit a signed, immutable evidence bundle (Part 11) with full traceability Everything is evidence (P4); the record is the validation artifact 3. Deterministic verification layer (the workhorse) # This layer is why the program works. Principle P5: the harness is the product. Each verifier is a deterministic function verify(artifact, context) → {PASS, FAIL(diagnostics)} that is reproducible (P7), versioned, and recorded. They compose: an artifact is admitted only if it survives the conjunction of all applicable verifiers for its class. Verifier class What it guarantees Critical-path position Determinism Build / compile Artifact is syntactically valid and integrates into the target Gate 0 — runs first; cheap, eliminates gross failures Fully deterministic Type systems / static typing Type-level contracts hold; whole classes of interface defects excluded Gate 0/1 — pre-test Fully deterministic Unit tests Specified behaviors hold on chosen inputs Gate 1 — core correctness Deterministic given seeded fixtures Property-based tests Invariants hold across generated input distributions, not just examples Gate 1/2 — depth beyond unit Deterministic with fixed seed; record seed in evidence Mutation testing The test suite itself is adequate (kills injected faults) — guards against vacuous green Gate 2 — meta-check on tests Deterministic; expensive, sampled or scheduled Differential / metamorphic testing New implementation matches a reference oracle, or obeys metamorphic relations when no oracle exists Gate 2 — oracle problem mitigation Deterministic Fuzzing No crashes/UB/memory-safety violations on adversarial inputs Gate 2/3 — robustness Deterministic per corpus + seed; time-boxed SAST / DAST / SCA No known-pattern vulnerabilities, runtime exposures, or vulnerable dependencies Gate 2 — security (ties to 07) Deterministic given ruleset + version pinning Schema / contract validation Interfaces, data, and API contracts conform; structured outputs are well-formed Gate 0/1 — fast structural guard Fully deterministic Formal methods / specification checks Critical properties provably hold (model checking, SMT, refinement) — reserved for Class C safety functions Gate 3 — strongest, narrowest Deterministic (proof or counterexample) Policy-as-code Org/regulatory rules enforced mechanically (licensing, banned APIs, segregation of duties, sign-off rules) Gate at every stage — governance Fully deterministic 3.1 How composition drives escape-rate down # Verifiers are arranged as a defense-in-depth conjunction , ordered cheap-to-expensive so that most bad artifacts are killed early at low cost (P6). If verifier i has independent miss-probability mᵢ (it lets a defect through), and a defect must evade all of them to escape, the residual escape probability is bounded by the product ∏ mᵢ — provided the verifiers are sufficiently independent (catch different defect families). Independence is the design objective: type checkers catch interface defects, property tests catch invariant violations, mutation testing catches weak tests , SAST catches vulnerability patterns. Correlated verifiers (e.g., two SAST tools sharing a ruleset) do not multiply their protection, and we do not claim they do (§6.1, §11). The Eval Owner (§8) maintains a documented defect-family-to-verifier coverage matrix so that the independence claim underlying the math is auditable rather than assumed. 4. Generation-time techniques that raise first-pass yield # These raise FPY and cut repair cost (an economic win, P6/08). They are not a substitute for the gate — they make the generator produce gate-passable artifacts more often. Technique Mechanism Yield effect Caveat Constrained / structured decoding Force output to a grammar/JSON schema/AST shape during generation Eliminates malformed-output failures outright Constrain form, not correctness — still must pass verifiers Spec-conditioned generation (BDD/Gherkin) Condition the model on executable acceptance criteria Aligns generation to the exact gate it must pass Spec quality bounds outcome quality Retrieval grounding Inject relevant internal code, standards, prior decisions into context Reduces hallucinated APIs and policy violations Retrieved context must itself be governed (07) N-sample self-consistency + verifier selection Generate N candidates; select by deterministic verifier outcome , not by model vote Raises probability ≥1 candidate passes; selection stays deterministic Cost ∝ N — budget per safety class (08) Test-first generation Agent writes a failing test from the spec, then code to satisfy it Forces an explicit, checkable success criterion Test must be reviewed so the agent doesn't write a weak/vacuous test (mutation testing guards this) Bounded repair loops Feed verifier diagnostics back; retry up to budget K Recovers near-misses cheaply Budget exhaustion → abstain/escalate , never force-pass Selection rule (load-bearing): when using N-sample self-consistency, the selector is the deterministic verifier suite , not an LLM majority vote. Self-consistency raises the chance a good candidate exists ; determinism decides which one is admitted . This keeps P2 intact even while exploiting probabilistic breadth. 5. Probabilistic evaluation (secondary, never sole gate) # Probabilistic evaluation has real value for coverage of qualities deterministic checks cannot express (clarity of a rationale, plausibility of a design narrative, reasonableness of a trajectory). It is admitted only under strict containment. Probabilistic signal What it assesses Permitted role Prohibited role Golden datasets + rubrics Behavior vs curated expected outcomes Trend/regression metric; gate only where ground truth is deterministic — LLM-as-judge Subjective quality, rationale adequacy Secondary signal ; escalation trigger; advisory on Class A only with human confirmation Sole gate for any Class B or Class C artifact Trajectory evaluation Did the agent take valid steps / legal tool calls in a sane order (via OpenTelemetry traces) Detects unsafe process even when output passes; flags for review Cannot admit on its own Behavioral drift detection Distribution shift in outputs/scores vs a baseline Monitoring + revalidation trigger (§9) Cannot gate a single release Gating rules (binding): For Class B and Class C , every admit decision must rest on a deterministic verdict. Probabilistic signals may only block (raise concern → escalate) or inform ; they may never admit (P2). LLM-as-judge is never the sole gate for B/C, full stop. Where used, the judge model, prompt, rubric, and version are recorded as part of the evidence and are themselves subject to drift monitoring. An LLM-as-judge disagreement with a deterministic verdict resolves in favor of the deterministic verdict , and the disagreement is logged for Eval Owner review. 6. The math of 99.9% # 6.1 Composing imperfect independent checks # Suppose a generated artifact is defective with prior probability p_def . It must pass n independent verifiers to be admitted; verifier i misses a defect with probability mᵢ . The probability a defect escapes the gate is bounded by: P(escape) ≤ p_def · ∏(i=1..n) mᵢ With three genuinely independent checks each missing 10% of defects ( mᵢ = 0.1 ) and a 20% defect prior, P(escape) ≤ 0.2 · 0.001 = 2e-4 — already below a 1e-3 ceiling. Add the bounded repair loop : each repair iteration re-subjects the artifact to the full conjunction, compounding the protection on the retained (repaired) artifacts. The headline number — release-gate correctness ≥ 99.9% — is therefore (b) precision achieved by composition , not by any single model or check. The independence caveat is the whole ballgame. The product rule holds only to the degree verifiers catch different defect families. We do not claim multiplicative protection for correlated checks; the Eval Owner's coverage matrix (§3.1) documents which mᵢ are credibly independent. Treating correlated checks as independent is the central statistical lie of "eval theater" (§11) and is explicitly disallowed. 6.2 Statistical acceptance by safety class # We never assert 99.9% from a point estimate. We require a lower confidence bound from a sample of sufficient size, scaled by safety class (risk-based, per CSA/GAMP 5). Safety class (IEC 62304) Acceptance criterion (placeholder) Sampling rigor Human sign-off Class A (no injury) Wilson/Clopper–Pearson lower 95% bound on gate precision ≥ 99.0% Standard golden suite Optional / sampled Class B (non-serious injury) Lower 95% bound ≥ 99.9%; escape-rate ≤ 1e-3 Expanded suite + property/mutation depth Required single reviewer Class C (death/serious injury) Lower 95–99% bound ≥ 99.9%; formal checks on safety functions; escape-rate ceiling set per risk file (14971) Maximum rigor; formal methods where feasible Required dual human control (P3) — non-negotiable To demonstrate a true rate ≥ 99.9% with a 95% lower confidence bound and zero observed failures , the rule-of-three approximation requires roughly n ≈ 3 / (1 − 0.999) ≈ 3000 independent passing items; any observed failure raises the required n substantially. Sample-size targets per suite are recorded with the eval suite version (§8). These are gate thresholds; they do not replace required human review. 6.3 Why Class C still needs human sign-off regardless of score # A 99.9% gate is a statistical statement about a population . A single Class C artifact may be the 1-in-1000 that the population statistics tolerate but a patient cannot. IEC 62304, ISO 14971, and CSA intended-use logic all require human judgment proportional to harm. Therefore dual human control on Class C is independent of the eval score — even a hypothetical 100% historical pass-rate does not waive it (P3). The eval score informs the reviewers; it never replaces them. 6.4 Abstention as a yield lever # A model that can say "I don't know" and escalate (per the fleet design in 04) converts a potential false pass into an honest escalation . Mathematically, abstention removes the hardest, least-certain items from the admitted population, which raises precision (b) and lowers escape-rate (c) at the cost of throughput (FPY) and human effort. We therefore treat a calibrated abstention-rate as a feature to tune , not a failure to suppress . Suppressing abstention to inflate FPY is an anti-pattern (§11). 7. Validation under FDA CSA / GAMP 5 / IEC 62304 # The agent and its harness are production/quality-system software that must be validated — not merely a developer convenience. We apply FDA Computer Software Assurance (risk-based assurance, intended use, recorded evidence), GAMP 5 (2nd ed) risk-based CSV, IEC 62304 V&V activities, and ISO/IEC 42001 for the AI management system. CSV/CSA element How it is satisfied here Intended-use statement Each agent/harness component has a documented intended use, operational boundaries, and the safety classes it may act on (drives risk-based rigor) Risk-based test rigor Test depth scales by IEC 62304 class (§6.2); CSA "unscripted vs scripted" assurance applied proportional to harm IQ analog Harness, model fleet, verifier toolchain deployed to K8s with pinned versions; installation verified and recorded (ties to 03) OQ analog Each verifier and gate exercised against golden suites; deterministic verdicts reproduced; thresholds demonstrated PQ analog End-to-end performance on representative change-sets at target precision/escape-rate; continuous in production (§9) Documented/recorded evidence Every gate emits a signed evidence bundle; nothing is asserted without a record (P4, Part 11) Traceability requirement → spec → code → test → eval → release , linked bidirectionally and stored immutably (01, 03) 7.1 Revalidation triggers # Validation is a state, not an event. Any of the following triggers risk-assessed revalidation, scoped by impact: Model change — new fine-tune, base-weight update, quantization, or decoding-config change (ties to 04). Harness change — new/updated verifier, gate-threshold change, repair-budget change, orchestration change (06). Eval-dataset change — new golden items, rubric change, contamination remediation (§8). Drift — behavioral/score drift past a control limit in production (§9). Regulatory/process change — change in intended use, standard revision, or risk-file update (14971). Each trigger, the affected scope, and the revalidation outcome are recorded; the traceability graph identifies the blast radius automatically. 8. Eval suites as controlled artifacts # Eval suites are controlled, validated artifacts with the same rigor as the software they judge — because under CSA they are quality-system software. Versioned & signed — each suite has a semantic version, content hash, and a Part 11 signature; gate runs record exactly which suite version produced the verdict (P7). Owned — an explicit Eval Owner role is accountable for suite integrity, the defect-family coverage matrix (§3.1), sample-size adequacy (§6.2), and contamination control. The Eval Owner is segregated from the model-training function (segregation of duties, policy-as-code enforced — §3, 07). Contamination control vs fine-tuning data — golden eval items are held out of, and continuously diffed against, fine-tuning corpora (ties to 04 ). Any leakage invalidates the affected results and triggers re-curation. Provenance and hashes of eval vs train sets are recorded so contamination is detectable , not merely asserted . Regression capture — every production escape becomes a new golden item (the closed loop in §2). This is the mechanism by which the system learns from failures and is the L5 "self-optimizing" behavior in 02 — but the new item enters the controlled suite through the same versioning/sign-off, never silently. 9. Continuous evaluation & monitoring in production # A gate that is only run pre-merge cannot detect post-deployment drift. Continuous evaluation closes the loop (PQ analog, §7). Capability What it does Ties to Live eval Periodically re-run golden suites against the current deployed fleet+harness; alert on regression §6.2 thresholds Drift detection Track score/behavior distributions against control limits; flag shift §5, §7.1 revalidation Auto-regression Re-run the captured-failure golden items on every model/harness change to prevent recurrence §8 MTTR Track mean-time-to-remediate gate regressions and production escapes; an operational KPI 09 Eval-cost budgeting Cost-aware scheduling of expensive checks (mutation, fuzzing, N-sample, formal); spend governed per cost-per-green-PR (P6) 08 Expensive deterministic checks are scheduled risk-proportionally: always-on for Class C critical paths, sampled or nightly for lower-risk surfaces, so assurance scales without unbounded GPU/compute cost (08). 10. Worked example: a Class B code change through every gate # A change-set modifies a Class B data-formatting routine. It is dispatched to a mid-tier (M) model with a constrained, spec-conditioned prompt. Step Action Evidence produced 0. Intent Requirement + Gherkin acceptance criteria retrieved and linked Trace edge: requirement → spec 1. Generate N=4 candidates via structured decoding; agent first writes a failing test 4 candidate diffs + 1 test, OTel trajectory trace 2. Build/type/schema Gate 0 run on all candidates; 1 fails compile, dropped Compile + type-check logs (deterministic) 3. Unit + property Surviving candidates run against unit + property suites; verifier selects the passing candidate Test results, property seeds recorded 4. Mutation Mutation run confirms the test suite kills injected faults (not vacuously green) Mutation score report 5. SAST/SCA Security scan clean; dependencies unchanged SAST/SCA report (07) 6. Eval Gate Golden suite for this surface meets Class B criterion (lower 95% bound ≥ 99.9%) Statistical acceptance record (§6.2) 7. LLM-as-judge Secondary rationale check — advisory only ; agrees, logged, does not gate Judge model+prompt+version, score (non-gating) 8. HITL Single qualified reviewer (Class B) approves with comments Signed human review record (Part 11) 9. Release record Signed evidence bundle assembled; full traceability sealed requirement→spec→code→test→eval→release graph Had any deterministic gate failed, the diagnostics would feed the bounded repair loop ; on budget exhaustion the item would abstain and escalate , never force-pass. Had this been Class C , step 8 would require dual human control and step 6 would add formal/specification checks regardless of the eval score (§6.3). 11. Anti-patterns # Anti-pattern Why it is dangerous Countermeasure Eval theater Impressive dashboards over weak/correlated checks; the 99.9% is unbacked Coverage matrix + independence audit (§3.1); mutation testing on the eval suite itself LLM-as-sole-gate A probabilistic judge admitting B/C artifacts violates P2; no reproducible verdict Hard rule: deterministic admit for B/C; judge is secondary/escalation only (§5) Training on the eval set Contamination inflates scores and hides true escape-rate Held-out signed suites, train/eval diffing, provenance hashes (§8, 04) Gaming first-pass yield by hiding repairs FPY treated as a target → incentive to suppress/mask repair iterations and abstentions FPY is an economic metric, never a gate; repairs and abstentions are recorded and audited (§1.1, §6.4) Treating correlated checks as independent Overstates the composition math; real escape-rate higher than claimed Document independence; only credit independent miss-probabilities (§6.1) Waiving Class C human control on a high score Statistics about a population do not protect the individual patient Dual control on C is independent of eval score (P3, §6.3) Summary # We do not sell a 99.9% model; we engineer a 99.9% system . Deterministic verifiers on the critical path, composed for independence and proven with statistical acceptance, produce the headline release-gate correctness and the low escape-rate. Probabilistic evaluation stays strictly secondary, abstention is a deliberate yield lever, human control scales with IEC 62304 risk, and every verdict is signed, traceable, reproducible evidence under CSA / GAMP 5 / Part 11. The eval system itself is validated, owned, and version-controlled — because in a regulated medical-device SDLC, the harness is the product (P5), and its evidence is the validation. ← Previous 04 · Model Strategy & Fine-Tuning Next → 06 · Agentic Workflows On this page 1. Reframing the 99.9%: what number are we actually defending? 2. The assurance pipeline 3. Deterministic verification layer (the workhorse) 4. Generation-time techniques that raise first-pass yield 5. Probabilistic evaluation (secondary, never sole gate) 6. The math of 99.9% 7. Validation under FDA CSA / GAMP 5 / IEC 62304 8. Eval suites as controlled artifacts 9. Continuous evaluation & monitoring in production 10. Worked example: a Class B code change through every gate 11. Anti-patterns
# 06 · Agentic Workflows · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/06-agentic-workflows.html
## 06 — Agentic Workflows #
06 · Agentic Workflows · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate 06 — Agentic Workflows # Figure C — Multi-Agent Workflow (gated & evidenced) · open SVG Part of Agentic-Native SDLC for Regulated Medical Device Engineering . Status: Working draft — May 2026. All numeric thresholds are placeholders pending calibration in 05-evaluation-and-validation.md . Scope: how individual and multi-agent workflows are designed, bounded, gated, and governed across the SDLC under IEC 62304, ISO 13485/QMSR, ISO 14971, FDA CSA, 21 CFR Part 11, and GAMP 5. This document operationalizes the seven principles into concrete, buildable workflows. It assumes the reference architecture in 03-reference-architecture.md (K8s, MCP tool plane, A2A, Argo Workflows, gVisor/Kata sandboxes, policy server, deterministic hooks, durable memory, OpenTelemetry trajectory tracing) and the model fleet tiers (S/M/L/V/E) in 04-model-strategy-and-finetuning.md . 1. Workflow design principles in a regulated setting # A workflow is a versioned, validated, change-controlled artifact — not an ad-hoc prompt chain. Every workflow MUST satisfy the following structural invariants before it is permitted to execute against a Class A/B/C codebase. # Invariant Mechanism Principle Failure mode it prevents WD-1 Bounded tasks Each agent step has an explicit input contract, output schema, max-iteration count, and acceptance predicate. No open-ended "do the thing" steps. P1, P5 Unbounded loops; scope creep WD-2 Deterministic gates between steps Inter-step transitions are mediated by deterministic verifiers (compilers, linters, test runners, schema validators, policy server). Probabilistic output never flows to the next step ungated. P2 Error propagation; non-reproducible state WD-3 HITL checkpoints by safety class Checkpoint density scales with IEC 62304 class; Class C requires dual human control . P3 Inappropriate autonomy on safety-critical code WD-4 Evidence capture per step Every step emits a signed evidence record (inputs, model+harness version, tool calls, gate results, diffs, approver) to the Part 11 evidence store. P4 Untraceable changes; audit gaps WD-5 Budget caps per loop Token + GPU + wall-clock budgets enforced per step and per workflow; breach → abstain/escalate, never silently truncate. P6 Runaway cost; "cost-per-green-PR" blowout WD-6 Abstention / escalation Agents emit a typed ABSTAIN(reason, evidence) rather than guessing when confidence/coverage falls below threshold. Escalation routes to HITL or a higher model tier. P1, P3 Confident-but-wrong output reaching a gate WD-7 Idempotency / rollback Every mutating action is idempotent and reversible: ephemeral sandbox branches, content-addressed artifacts, transactional commits, one-command rollback. P7 Partial/irreversible damage to mainline The harness is the product (P5). Workflows are expressed declaratively as Argo Workflow templates whose steps invoke agents through the harness; the LLM is a swappable component (see §3, Agent = Model + Harness ). The cost-per-green-PR metric (P6) is the primary economic objective function for every workflow (see 08-token-and-gpu-economics.md ). flowchart LR G[Generate
probabilistic] --> V[Verify
deterministic gate] V -->|fail| R[Repair
bounded loop] R --> V V -->|pass| GA[Gate
policy + eval + HITL] GA -->|reject| R GA -->|approve| M[(Mainline)] V -->|budget/iter breach| ESC[Abstain → Escalate] R -->|budget/iter breach| ESC classDef det fill:#e8f0fe,stroke:#1a73e8; classDef prob fill:#fce8e6,stroke:#d93025; class V,GA det; class G,R prob; This Generate → Verify → Repair → Gate loop is the atomic unit of every workflow in this document and is the realization of the 99.9% release-gate correctness system property (P1): individual model calls are unreliable; the loop is engineered to be reliable. 2. Operating modes: Conductor vs Orchestrator # Two operating modes cover the spectrum from synchronous developer assistance to autonomous async batch work. A given task is routed to a mode by the policy server based on safety class, change size, and required latency. Dimension Conductor (real-time, in-IDE) Orchestrator (async multi-agent) Interaction model Synchronous; human in the loop continuously Asynchronous; human at defined checkpoints Latency target Sub-second to seconds Minutes to hours Primary surface IDE / editor / terminal Argo Workflows, A2A mesh, CI Concurrency Single active agent, human-paced Many sub-agents in parallel (Planner/Coder/Test/…) Control granularity Per-edit, per-tool-call (human approves inline) Per-stage gate + HITL checkpoints Typical model tier S/M (low latency), V for visual context M/L (capable), E for hard reasoning; routed per step Tooling MCP tools scoped to open repo; pre-tool/post-edit hooks Full MCP plane; sandbox per sub-agent; policy server inline Memory Session-scoped; durable per-developer Durable shared workflow memory + per-agent scratch Evidence Lightweight (suggestion accepted/rejected) Full per-step signed evidence chain Best for Conductor: pair-style implementation, refactor-in-place, explain, local test authoring Orchestrator: feature delivery, coverage expansion, migrations, repo-watching, batch fixes Autonomy ceiling Class A/B with human approving each apply Up to L4 within validated bounds; Class C still dual-control ASMM-Med fit L1–L2 L2–L5 A workflow may hand off between modes: a developer in Conductor mode dispatches a bounded task to the Orchestrator ("expand coverage on pump_controller "), reviews the resulting PR, and pulls it back into Conductor for final touch-ups. All handoffs are A2A messages with attached context manifests. sequenceDiagram participant Dev as Developer (IDE) participant C as Conductor Agent participant O as Orchestrator (Argo) participant Sub as Sub-agents (A2A) Dev->>C: "implement story PUMP-142" C->>C: scope check (safety class B) C->>O: dispatch bounded task + context manifest O->>Sub: fan-out Planner/Coder/Test/Review Sub-->>O: gated PR + evidence O-->>Dev: HITL checkpoint (PR review) Dev->>C: pull back for local refinement 3. Agent anatomy — Agent = Model + Harness # The platform's first axiom: an agent is not a model . An agent is a model wrapped in a deterministic harness that supplies capability, constraint, and evidence. The model is the only probabilistic component; everything else is engineered software under change control. flowchart TB subgraph Agent direction TB M[Model — fleet tier S/M/L/V/E
self-hosted, fine-tuned, open-weight] subgraph Harness I[Instructions / rule files
AGENTS.md, specs/, BDD/Gherkin] T[MCP tool plane
typed, permissioned tools] S[Sandbox
gVisor/Kata, ephemeral, egress-deny] OR[Orchestration
Argo steps + A2A handoffs] H[Hooks / guardrails
pre-tool, post-edit, pre-commit] ME[Memory
durable sessions + scoped scratch] OB[Observability
OpenTelemetry trajectory tracing] PS[Policy server
structural + semantic gating] end end M <--> I M <--> T T --> PS T --> S M --> H M --> ME Agent --> OB Component Role Determinism Cross-ref Model (S/M/L/V/E) Generation, reasoning, repair proposals Probabilistic 04 Instructions / rule files AGENTS.md , specs/ , BDD/Gherkin define behavior, conventions, constraints Deterministic source-of-truth 01 MCP tools Typed, permissioned actions (read/edit/test/build/query) Deterministic interface 03 Sandbox gVisor/Kata ephemeral env; egress-deny; per-task Deterministic isolation 07 Orchestration Argo step graph + A2A handoffs Deterministic 03 Hooks / guardrails Lifecycle interceptors (pre-tool, post-edit, pre-commit) Deterministic §9 Memory Durable session/workflow memory + scoped scratch Deterministic store 03 Observability OTel trajectory spans for every tool call and gate Deterministic 05 Policy server Structural + semantic gating, intercepts every tool call Deterministic 07 The policy server sits between the model and every tool : a model may propose a write_file or run_command , but the action only executes if it passes structural rules (path allowlist, diff size, no secret egress) and semantic rules (change consistent with the active spec). This is how "determinism wraps probabilism" (P2) is enforced at the tool boundary. 4. Multi-agent decomposition pattern # The Orchestrator decomposes a unit of work into specialized sub-agents connected by A2A handoffs. Each handoff crosses a deterministic gate. Humans sign at safety-class-appropriate checkpoints. flowchart LR H[Human request / story] --> PL[Planner] PL -->|A2A: plan| SP[Spec / Story agent] SP -->|A2A: spec + Gherkin| G1{Spec gate
policy + eval} G1 -->|approve| HC1[[HITL: spec sign-off]] HC1 --> CO[Coder] CO -->|A2A: diff| G2{Build/lint gate} G2 --> TE[Test agent] TE -->|A2A: tests+coverage| G3{Test+coverage gate} G3 --> RV[Review agent] RV -->|A2A: findings| G4{Review gate
policy + eval} G4 --> HC2[[HITL: PR approval
Class C = dual control]] HC2 --> IN[Integrator] IN -->|A2A: merge req| G5{Release gate ≥99.9%} G5 --> M[(Mainline / deploy)] G1 -.reject.-> SP G2 -.reject.-> CO G3 -.reject.-> CO G4 -.reject.-> CO Policy server is invoked inside every gate (G1–G5) and on every tool call within each sub-agent. Eval gates (deterministic evaluation harness, 05 ) sit at G1, G4, G5. Humans sign at HC1 (spec) and HC2 (PR); Class C requires two independent approvers at HC2 (P3). A2A messages are content-addressed and carry the evidence manifest forward so the Integrator can assemble a complete Part 11 record. 5. Per-SDLC-phase workflows # Sub-template used for every phase: Goal · Agent role · Inputs/context · Tools · Deterministic gates · HITL checkpoint · Evidence · IEC 62304 activity · Model tier. 5.1 Requirements / planning # Field Value Goal Convert stakeholder intent into structured, testable requirements + acceptance criteria Agent role Spec/Story agent (decompose, normalize, detect ambiguity/conflict) Inputs / context Intake notes, existing specs/ , risk file (ISO 14971), product requirements Tools requirements_db , traceability_query , risk_register , MCP doc retrieval Deterministic gates Schema validation of requirement records; traceability completeness check; duplicate/conflict linter HITL checkpoint Requirements review board sign-off (mandatory all classes) Evidence Versioned requirements set, ambiguity report, trace links, approver record IEC 62304 5.2 Software requirements analysis Model tier L (reasoning over ambiguity); E for high-risk Class C decomposition 5.2 Design / architecture # Field Value Goal Produce software architecture + detailed design consistent with requirements and risk controls Agent role Design agent (architecture proposal, interface contracts, design rationale) Inputs / context Approved requirements, architecture standards, risk controls, existing design docs Tools architecture_model , diagram_gen , interface_registry , dependency graph query Deterministic gates Interface contract validation; architecture rule checks; risk-control coverage check HITL checkpoint Design review (mandatory); Class C requires safety reviewer Evidence Design records, interface specs, design-to-requirement trace, decision log IEC 62304 5.3 Software architectural design; 5.4 detailed design Model tier L; V for diagram/visual artifacts 5.3 Implementation # Field Value Goal Implement units to satisfy design + spec, passing all deterministic gates Agent role Coder (generate diff in sandbox, self-repair against gates) Inputs / context Approved design, specs/ , AGENTS.md , target files, dependency graph Tools read_file , write_file , run_build , run_lint , run_unit_tests , static_analyzer Deterministic gates Compile/build, lint, static analysis (SAST), unit tests, diff-size policy HITL checkpoint Diff review on PR; Class B/C never auto-merged without diff review Evidence Diff, build/test logs, static-analysis report, model+harness version IEC 62304 5.5 Software unit implementation and verification Model tier M (default); L for complex units; routed by complexity stateDiagram-v2 [*] --> Plan Plan --> Generate: bounded task Generate --> Verify Verify --> Repair: gate fail Repair --> Verify Verify --> Abstain: iter/budget breach Verify --> PR: all gates green Abstain --> Escalate PR --> [*]: diff review (HITL) 5.4 Test / QA # Field Value Goal Author/extend verification tests; achieve coverage + behavioral targets Agent role Test agent (generate tests, mutation-check, close coverage gaps) Inputs / context Code under test, requirements/Gherkin, existing tests, coverage baseline Tools run_tests , coverage_tool , mutation_tester , requirements_trace Deterministic gates Tests pass; coverage ≥ threshold (placeholder); mutation score ≥ threshold; no flaky/quarantine regressions HITL checkpoint QA lead review of new test suite; Class C verifies requirement-to-test trace Evidence Test suite diff, coverage delta, mutation report, requirement-test trace matrix IEC 62304 5.5 unit verification; 5.6 integration testing; 5.7 system testing Model tier M; L for hard property/edge-case synthesis flowchart LR C[Code under test] --> TA[Test agent] TA --> GEN[Generate tests] GEN --> R{Run tests} R -->|fail to author| TA R -->|pass| COV{Coverage ≥ θ?} COV -->|no| TA COV -->|yes| MUT{Mutation ≥ θ?} MUT -->|no| TA MUT -->|yes| TR{Req-trace complete?} TR -->|yes| QA[[HITL: QA review]] 5.5 Code review # Field Value Goal Detect defects, spec deviations, risk-control violations before merge Agent role Review agent (semantic diff review, standards + risk checks) Inputs / context PR diff, spec, design, coding standards, risk controls, prior findings Tools diff_view , policy_check , standards_linter , trace_query , sec_review Deterministic gates Policy server semantic gate; standards lint; no unresolved high-severity findings HITL checkpoint Human reviewer approves; Class C dual control; conditional-LGTM only where permitted (§7) Evidence Review findings, resolution log, approver(s), gate results IEC 62304 5.5/5.6 verification; supports 9 problem resolution Model tier L (judgment); M for routine diffs sequenceDiagram participant PR as Pull Request participant RA as Review Agent participant PS as Policy Server participant Ev as Eval Gate participant Hu as Human Reviewer(s) PR->>RA: diff + context RA->>PS: semantic + structural check PS-->>RA: pass/findings RA->>Ev: deterministic review eval Ev-->>RA: score ≥ θ RA->>Hu: findings + recommendation Hu-->>PR: approve (dual for Class C) 5.6 Deployment # Field Value Goal Promote validated build through release gate to target environment Agent role Integrator/Release agent (assemble release record, drive pipeline) Inputs / context Approved PR, full evidence chain, release checklist, change-control record Tools argo_pipeline , release_gate , evidence_store , signing_service , deploy Deterministic gates Release gate ≥99.9% correctness ; evidence completeness; signed approvals present HITL checkpoint Release authority sign-off (mandatory); Class C dual authority Evidence Signed release record, DHF/Part 11 package, deploy manifest, rollback plan IEC 62304 5.8 software release Model tier S/M (orchestration, not generation) 5.7 Maintenance / legacy modernization # Field Value Goal Remediate defects; modernize legacy code with behavior preservation Agent role Maintenance/Migration sub-agent pipeline (graph-native understanding) Inputs / context Defect report or migration scope, code graph, characterization tests, risk file Tools code_graph_query , characterization_tests , run_tests , diff_view , equivalence_check Deterministic gates Characterization tests pass pre/post; behavioral equivalence; coverage maintained HITL checkpoint Change review; Class C dual control; CAPA linkage for defects Evidence Defect-to-fix trace, before/after behavior proof, migration manifest IEC 62304 6 software maintenance; 9 problem resolution Model tier L (analysis) + M (bulk edits); E for hard equivalence reasoning 6. Concrete worked example workflows # 6.1 Feature implementation on a Class B module # Scenario: Story PUMP-142 adds a configurable alarm threshold to infusion_rate_monitor (Class B). Step Mode/agent Action Gate Evidence 1 Orchestrator / Spec Decompose story → spec + Gherkin acceptance Schema + trace Spec record, trace links 2 HITL Spec sign-off (single approver, Class B) — Approver record 3 Coder (M) Implement in sandbox; self-repair vs build/lint/unit Compile, lint, SAST, unit Diff, logs 4 Test (M) Add tests for new branches; close coverage Coverage ≥ θ, mutation ≥ θ Coverage delta, mutation report 5 Review (L) Semantic review vs spec + risk controls Policy semantic gate Findings + resolutions 6 HITL PR diff review (mandatory, single for Class B) — Approval 7 Integrator (S) Release gate, assemble record ≥99.9% release gate Signed release record Budget cap: workflow aborts and escalates if token+GPU spend exceeds the per-PR cap (P6, 08 ). No auto-merge — Class B requires human diff review (§10 anti-pattern). 6.2 AI-generated test-coverage expansion # Scenario: Raise coverage on dosage_calculator from 71% to ≥ target without changing behavior. flowchart LR BL[Baseline coverage] --> GAP[Coverage-gap analysis
code graph] GAP --> TA[Test agent: synth tests] TA --> RUN{Tests green?} RUN -->|no| TA RUN -->|yes| FLK{Flaky? quarantine check} FLK -->|stable| MUT{Mutation ≥ θ} MUT -->|yes| BEH{No behavior change
vs baseline} BEH -->|confirmed| QA[[HITL: QA lead]] BEH -->|drift| TA Key controls: tests must be additive and behavior-preserving — any test that would have failed against unchanged production code is flagged as a latent defect and escalated rather than silently "fixed." Mutation testing guards against vacuous tests. Evidence: coverage delta, mutation score, requirement-test trace. 6.3 Bug fix in forensic mode (failing-test-first, evidence prompting) # Scenario: Field complaint → defect DEF-908 in battery_health_estimator (Class C). Forensic mode enforces reproduce-before-repair. stateDiagram-v2 [*] --> Reproduce Reproduce --> WriteFailingTest: capture defect as test WriteFailingTest --> ConfirmRed: test fails on current code ConfirmRed --> RootCause: evidence-prompted analysis RootCause --> Fix: bounded minimal diff Fix --> Green: failing test now passes Green --> Regression: full suite + characterization Regression --> DualReview: Class C dual control DualReview --> CAPA: link to problem resolution CAPA --> [*] Failing-test-first: the defect is encoded as a test that is red before any fix; this becomes permanent regression evidence. Evidence prompting: the root-cause step is required to cite specific code-graph nodes, traces, and the failing assertion — no unsupported hypotheses. Class C: dual human control at review; fix is linked to CAPA / IEC 62304 §9 problem resolution. Minimal-diff policy: the repair loop is bounded to the smallest change that turns the test green; scope expansion triggers escalation. 6.4 Legacy modernization / framework migration at scale # Scenario: Migrate ~400 modules from a deprecated UI framework to the supported one, behavior-preserving, across Class A/B code. flowchart TB SC[Scope intake] --> GRAPH[Graph-native code understanding
build dependency + call graph] GRAPH --> CHAR[Characterization test harvest
pin current behavior] CHAR --> PART[Partition into wave batches
by risk + coupling] PART --> FAN[Fan-out sub-agent pipeline] subgraph perModule[Per-module pipeline] MIG[Migration coder] --> EQ{Behavioral equivalence} EQ -->|fail| MIG EQ -->|pass| RV2[Review agent] end FAN --> perModule RV2 --> HITL[[HITL: batch review]] HITL --> INT[Integrator: staged merge] INT --> M[(Mainline)] Graph-native understanding: the migration is planned over the actual code/dependency graph, not file-by-file, so coupling and ordering are respected. Characterization tests pin pre-migration behavior; the equivalence gate is the deterministic guarantee of behavior preservation. Batching by risk: Class A modules may use lighter checkpoints; Class B retain mandatory diff review; any Class C in scope keeps dual control. Cost discipline: per-module budget caps and tier routing (M for mechanical edits, L for complex ones) keep cost-per-green-PR within bounds at scale. 7. Human-in-the-loop design # HITL is the mechanism for risk-proportional autonomy (P3). Checkpoint placement is a function of IEC 62304 safety class. Safety class Spec sign-off Diff review Release approval Auto-merge on green Class A Required (may batch) Required for non-trivial; conditional-LGTM permitted for low-risk additive change Single Permitted only where policy explicitly allows (additive, non-safety, fully gated) Class B Required Mandatory , per-PR diff review Single Not permitted Class C Required Mandatory, dual independent reviewers Dual authority Never Design measures to keep HITL effective without inducing approval fatigue: Dual-control for Class C: two independent qualified humans; the system enforces approver disjointness (no self-approval, no single person satisfying both). Batching: related low-risk approvals are grouped into a single review surface with shared context, reducing context-switch cost while preserving per-item evidence. Digital quiet hours: the Orchestrator respects configured quiet windows; checkpoints queued during quiet hours are surfaced at the next active window, never auto-escalated to auto-approval. Conditional-LGTM (merge on green): permitted only for Class A, additive, non-safety-critical changes that pass the full deterministic gate chain and the release gate; explicitly disabled for Class B/C. Every conditional-LGTM merge still produces a full evidence record and is sampled into the audit pipeline. Escalation routing: abstentions and budget breaches route to a named human queue, not into a silent retry storm. 8. Continuous code-review and repo-watcher agents # Continuous agents observe the repository and act on events (PR opened, push, scheduled scan). They are deployed in three tiers by integration depth. Tier Deployment Trigger Catches Evidence posting T1 — managed-equivalent Self-hosted equivalent of a managed review bot PR opened/updated Style, obvious bugs, lint/standards drift, secret leaks Inline PR comments + evidence record T2 — hybrid CI-triggered Review agent invoked from CI pipeline CI stage on PR/push Above + spec deviation, coverage/mutation regressions, risk-control violations CI check + structured findings to evidence store T3 — custom A2A Full Orchestrator sub-agent in the A2A mesh Event or schedule (repo-watcher) Above + cross-module/arch drift, dependency risk, traceability gaps; can open remediation PRs Signed findings, optional auto-PR, OTel trajectory All tiers route every proposed action through the policy server and post evidence to the Part 11 store. Repo-watcher agents (T3) operate under strict budget caps and an action allowlist; they may propose remediation PRs but never merge to Class B/C without the §7 HITL path. Findings are linked to requirements/risk items for traceability. 9. Workflow governance # A workflow is a validated software item in its own right and is managed under the QMS (ISO 13485/QMSR, GAMP 5). Versioning. Each workflow (Argo template + agent configs + rule files + tool permissions) is a content-addressed, semantically-versioned artifact. The exact model+harness versions used by a run are pinned and recorded (P7 reproducibility). Validation (tie to 05 ). Before promotion, a workflow is validated against the deterministic evaluation harness: golden task sets, gate-correctness measurement toward the ≥99.9% release-gate property, abstention calibration, and cost envelope. Validation evidence is part of the workflow's release record. CSA-aligned, risk-based validation depth scales with the highest safety class the workflow may touch. Change control. Workflow changes follow the same change-control and approval path as code: proposed diff → review → eval gate → approval → versioned release. A workflow change that alters gate behavior or autonomy level requires re-validation. flowchart LR WF[Workflow change proposal] --> RV[Review] RV --> EV[Eval harness validation\n→ 05] EV -->|meets ≥99.9% + cost| AP[Approval / change control] EV -->|fails| WF AP --> REL[Versioned workflow release] REL --> REG[(Registry — pinned)] Cost guardrails in-loop (tie to 08 ). Per-step and per-workflow token/GPU/wall-clock budgets are enforced at runtime by the orchestrator and hooks. Breach behavior is deterministic: pause → abstain → escalate. Cost telemetry feeds the cost-per-green-PR metric (P6) and the economics dashboards. Failure handling and rollback. Every workflow defines: (a) idempotent steps over ephemeral sandbox branches; (b) a one-command rollback to the last known-good mainline state; (c) a quarantine path for flaky/unstable artifacts; (d) escalation to HITL on repeated gate failure. No partial state ever reaches mainline — integration is transactional behind the release gate. Failure Detection Response Gate fail (transient) Verifier Bounded repair loop Iteration/budget breach Orchestrator counter Abstain → escalate to HITL Non-reproducible result Eval/replay Pin freeze; block promotion; investigate Bad merge slipped Release-gate audit / repo-watcher Automated rollback + CAPA 10. Anti-patterns # Anti-pattern Why it is dangerous Mitigation in this platform Unbounded loops Runaway cost; non-terminating agents; eroded determinism WD-1/WD-5 hard iteration + budget caps; abstain on breach Multi-file autonomous edits without diff review on Class B/C Unreviewed safety-relevant change reaches mainline §7 mandatory diff review; auto-merge disabled for B/C; policy server diff-size + path gates Agent-to-agent error amplification One sub-agent's hallucination becomes the next's "fact"; compounding error across A2A WD-2 deterministic gate at every handoff; no probabilistic output flows ungated; evidence carries provenance Context fragmentation Sub-agents work from inconsistent/partial context; divergent assumptions Single source-of-truth ( specs/ , AGENTS.md ); content-addressed context manifests on every A2A handoff; durable shared memory Confident wrong abstention-suppression Agent guesses instead of abstaining WD-6 typed ABSTAIN ; calibration validated in 05 Mode misuse (Conductor for batch) Latency/cost mismatch; weak evidence Policy-server routing by class/size/latency (§2) Cross-references # 01-requirements.md · 02-maturity-model.md · 03-reference-architecture.md · 04-model-strategy-and-finetuning.md · 05-evaluation-and-validation.md · 07-security-and-compliance.md · 08-token-and-gpu-economics.md · 09-adoption-roadmap.md ← Previous 05 · Evaluation & Validation Next → 07 · Security & Compliance On this page 1. Workflow design principles in a regulated setting 2. Operating modes: Conductor vs Orchestrator 3. Agent anatomy — Agent = Model + Harness 4. Multi-agent decomposition pattern 5. Per-SDLC-phase workflows 6. Concrete worked example workflows 7. Human-in-the-loop design 8. Continuous code-review and repo-watcher agents 9. Workflow governance 10. Anti-patterns
# 07 · Security & Compliance · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/07-security-and-compliance.html
## Security & Compliance — Agentic SDLC for Regulated Medical Device Engineering #
07 · Security & Compliance · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate Security & Compliance — Agentic SDLC for Regulated Medical Device Engineering # Audience: CISO and security architecture, Quality/Regulatory (QA/RA), MLOps/Platform, and an FDA / Notified Body auditor. Scope: The security and compliance control set for AI agents that build, test, document, and maintain regulated medical-device software, running Kubernetes-native with self-hosted, fine-tuned open-weight models only (no Claude/OpenAI/Gemini SaaS APIs — hard-blocked at the network layer). Date: May 2026. All numeric thresholds in this document are placeholders pending organizational calibration. Companion docs: 01-requirements · 02-maturity-model · 03-reference-architecture · 04-model-strategy-and-finetuning · 05-evaluation-and-validation · 06-agentic-workflows · 08-token-and-gpu-economics · 09-adoption-roadmap 0. Purpose and framing # This document is the security and compliance register for the agentic SDLC. It assumes the seven principles defined in 02-maturity-model — in particular P3 (risk-proportional autonomy by IEC 62304 class), P4 (everything an agent does is evidence, 21 CFR Part 11-grade), P5 (the harness is the product), and P7 (self-hosted, sovereign, reproducible). The governing security thesis: a probabilistic agent is an untrusted actor inside a zero-trust system. We do not trust the model to behave; we constrain what any model can do via deterministic structural controls, sandboxing, and signed supply-chain assurance, and we make every action attributable and replayable so it survives an audit or a recall investigation. Two regulatory and security framings run in parallel and must never be conflated (see §9): Track A — AI that BUILDS the device: production/Quality-System tooling. Validated under FDA Computer Software Assurance (CSA) , ISO 13485/QMSR, 21 CFR Part 11. This framework's primary focus. Track B — AI shipped INSIDE the device: SaMD / AI-enabled device function, submission-bearing, governed by EU AI Act high-risk obligations, FDA premarket cybersecurity, and a Predetermined Change Control Plan (PCCP). Threat frameworks of record: OWASP Top 10 for LLM Applications (2025) and MITRE ATLAS . Standards anchors: 21 CFR Part 11, IEC 62304, ISO 13485/QMSR, ISO 14971, ISO/IEC 42001 (AI management system), IEC 62443 (cyber for medical/industrial), FDA premarket cybersecurity guidance, FD&C Act §524B (SBOM + vulnerability management), EU MDR + EU AI Act. 1. Threat model for agentic dev in a regulated org # 1.1 Assets under protection # Asset Why it matters Loss class Source IP (regulated product source, DHF, algorithms) Core competitive + regulated artifact Confidentiality, Integrity PHI/PII May appear in test fixtures, bug repros, logs, support data Confidentiality (HIPAA/GDPR), regulatory breach Model weights + adapters (fleet S/M/L/V/E) Fine-tuned on proprietary corpora; theft = IP loss + cloneable behavior Confidentiality, Integrity Signing keys (Sigstore/cosign, Vault transit, SPIRE CA) Root of all provenance trust; compromise forges everything downstream Integrity (catastrophic) The regulated product itself A malicious or erroneous agent commit can injure patients Safety, Integrity Audit/evidence store (WORM logs) The Part 11 record of truth; if tamperable, nothing is defensible Integrity, Non-repudiation Eval gold sets (see 05 ) Poisoning the eval = silently lowering the release gate Integrity 1.2 Adversaries # Adversary Capability Primary objective External attacker Network probing, supply-chain injection, poisoned public repos/docs Exfiltration, foothold, weight theft Malicious insider Authenticated dev/operator access IP theft, sabotage, gate bypass Compromised dependency / model Trojaned open-weight base model, poisoned dataset, malicious package Backdoor in shipped product The agent itself (untrusted-by-design) Whatever tools it is granted; subject to injection/poisoning Unintended/rogue action, data exfil Negligent user Over-broad prompts, pasting PHI, approving without review Accidental leak, gate erosion 1.3 Agent-specific attack surface → framework mapping # Threat Description in our context OWASP LLM Top 10 MITRE ATLAS Prompt injection Malicious instructions in an issue, code comment, PR body, or doc the agent reads LLM01 AML.T0051 (LLM Prompt Injection) Context poisoning Tainted retrieval source / repo seeds the agent's working context LLM01, LLM08 AML.T0070 (RAG poisoning) Tool misuse Agent invokes a granted tool with harmful args (e.g., mass email, prod write) LLM06 (Excessive Agency) AML.T0053 (LLM Plugin Compromise) Data exfiltration Source/PHI/weights leak via tool output, egress, or model channel LLM02 (Sensitive Info Disclosure) AML.T0024, AML.T0057 Model supply-chain Trojaned base weights / poisoned adapter / malicious dataset LLM03 (Supply Chain) AML.T0010 (ML Supply Chain Compromise) Rogue autonomous action High-autonomy agent takes an unbounded irreversible action LLM06 AML.T0048 (External Harms) Training-data poisoning Corrupted fine-tune corpus embeds a backdoor or skews behavior LLM04 (Data/Model Poisoning) AML.T0020 (Poison Training Data) Insecure output handling Unsanitized agent output executed downstream (e.g., shell, SQL) LLM05 AML.T0050 System prompt / config leak Harness policy/secrets exposed via model LLM07 AML.T0056 1.4 Trust-boundary diagram # flowchart TB subgraph EXT["UNTRUSTED — external"] PUB[Public repos / docs / packages] SAAS[(External LLM APIs — HARD BLOCKED)] ADV([External attacker]) end subgraph EDGE["TB1: Org perimeter — egress deny-by-default"] EG{{Egress allow-list proxy}} SAAS -. "BLOCKED at L3/L7" .-x EG end subgraph MESH["TB2: K8s + Istio mTLS mesh — SPIFFE/SPIRE identity"] direction TB subgraph CTRL["Control plane (trusted)"] POL[Policy Server\nstructural + semantic gating] OPA[OPA / Gatekeeper admission] VAULT[(HashiCorp Vault)] AUD[(WORM audit / evidence store)] REG[(Signed model + artifact registry)] end subgraph SBX["TB3: Agent sandbox — ephemeral, low-priv"] AGENT([Dev agent\ngVisor/Kata]) TOOLS[Tools: git, build, test, browser, term] end INF[[Self-hosted inference\nfleet S/M/L/V/E]] end subgraph REG_ASSETS["TB4: Regulated assets (highest trust)"] SRC[(Source IP / DHF)] PROD[(Regulated product / release branch)] KEYS[(Signing keys / SPIRE CA)] end PUB --> EG --> AGENT AGENT -- "every tool call" --> POL POL -- allow/deny/sanitize --> TOOLS POL --> AUD AGENT <-->|mTLS| INF TOOLS -->|Vault-brokered, short-TTL| VAULT AGENT -. "no default write" .-x PROD POL -- "Class C: dual human control" --> PROD REG --> OPA --> SBX ADV -.-x EDGE Trust boundaries: TB1 perimeter (egress control), TB2 mesh (identity + mTLS), TB3 sandbox (blast-radius containment), TB4 regulated assets (signed, dual-controlled). An agent never crosses from TB3 to TB4 except through the Policy Server and (for Class B/C) human authorization. 2. Zero-trust architecture # Zero trust here means: no implicit trust by network location; every workload authenticates; every call is authorized; deny by default. Control Implementation What it enforces Workload identity SPIFFE/SPIRE — every agent, tool, model server gets a SPIFFE ID (SVID), attested at startup, short-TTL, auto-rotated No shared service accounts; every action attributable to a cryptographic identity (feeds P4 attribution) mTLS everywhere Istio service mesh; PeerAuthentication: STRICT mesh-wide No cleartext intra-mesh traffic; no spoofed peers Least-privilege RBAC K8s RBAC + AuthorizationPolicy keyed on SPIFFE ID; agents get only the namespaces/tools their role requires An agent role cannot reach services outside its task scope Network deny-by-default Default-deny NetworkPolicy (Cilium); explicit allow per workload pair Lateral movement blocked; sandbox cannot reach the audit store directly Egress allow-list L3/L7 egress proxy; allow-list of approved internal endpoints only Exfiltration channel closed Hard block on external LLM endpoints Egress proxy + DNS sinkhole deny api.openai.com , *.anthropic.com , generativelanguage.googleapis.com , etc.; alert + auto-quarantine on attempt Enforces P7 sovereignty; an injected agent cannot phone an external model out The external-LLM block is both a control and a detector : any attempt is treated as a potential prompt-injection/exfil indicator and raises a security incident (§11). 3. The Policy Server in depth # The Policy Server is the deterministic chokepoint through which every tool call passes before execution. It realizes principle P2 (determinism wraps probabilism) at the action boundary. It has two stages. 3.1 Structural gating (deterministic, policy-as-code) # Rules are pure functions of (role, environment, tool, args, safety_class) — no model in the loop, fully testable, version-controlled, signed. This is the authoritative, non-bypassable layer. yaml # policies.yaml (illustrative — versioned, cosign-signed, loaded read-only) version: 2026.05 defaults: effect: deny # deny-by-default roles: dev-agent-classA: allow: - tool: git.read - tool: build.run - tool: test.run - tool: pr.propose # propose only, never merge - tool: browser.fetch constraints: { url_allowlist: ["internal-docs", "approved-mirror"] } deny: - tool: email.send # structurally forbidden for any dev agent - tool: prod.write dev-agent-classC: allow: - tool: git.read - tool: test.run - tool: pr.propose require: - control: dual_human_authorization # see §10 matrix deny: - tool: "*.write" - tool: email.send - tool: browser.fetch # no external context for Class C work environments: prod: deny: - tool: "*.write" unless_role: ["release-bot-signed"] and_control: dual_human_authorization guards: - id: secret-egress match: { args_contains_secret: true } effect: deny - id: pii-in-args match: { semantic.pii_detected: true } # result from semantic stage effect: sanitize # mask then allow, or escalate Representative invariants (canonical examples reused across docs): role email.send is never granted to a dev agent; prod.write requires a signed release role and dual human authorization; Class C work forbids external browser context. 3.2 Semantic gating (intent / PII check before tool execution) # A secondary self-hosted LLM (a small fleet-tier model, see 04 ) inspects the intent and content of a tool call that passed structural rules but warrants judgment — e.g., does this pr.propose body or browser.fetch argument carry PHI/PII, exfiltration intent, or signs of injection? Semantic gating is advisory and escalating , never a sole gate (P2): it can downgrade (sanitize) or escalate (require human), but it can never grant something structural rules denied. 3.3 Interception pseudo-flow # text on tool_call(agent_id, role, env, tool, args): record = open_evidence_span(agent_id, role, tool, args_hash) # P4 # STAGE 1 — structural (deterministic, authoritative) s = structural_eval(role, env, tool, args, safety_class) if s == DENY: emit_evidence(record, decision=DENY, stage=structural); return BLOCKED # STAGE 2 — semantic (judgment; PII/intent/injection) sem = semantic_model.assess(tool, args, retrieval_provenance) if sem.pii or sem.exfil_intent or sem.injection_signal: if policy.allows_sanitize(tool): args = mask_placeholders(args) # [[VAR]] injection, §5 emit_evidence(record, decision=SANITIZE, findings=sem) else: emit_evidence(record, decision=ESCALATE, findings=sem) return REQUIRE_HUMAN(record) # STAGE 3 — control requirements (autonomy matrix, §10) if requires_dual_control(role, env, safety_class): emit_evidence(record, decision=PENDING_DUAL_CONTROL) return REQUIRE_HUMAN(record, control=DUAL) emit_evidence(record, decision=ALLOW) return EXECUTE(tool, args) Every branch emits evidence : input hash, structural verdict, semantic findings, sanitization diff, human-decision pointer, model+policy versions. This record is the Part 11 artifact (§7). 4. Agent sandboxing & blast-radius control # The agent runs untrusted-by-design; the sandbox guarantees that even a fully-compromised agent has a small, recoverable blast radius. Control Implementation Ephemeral runtime One agent run = one fresh ephemeral namespace + pod, torn down on completion; no persistence across runs Kernel-isolated sandbox gVisor (default) or Kata Containers (stronger isolation for V/E tiers or external-content tasks) — syscall surface contained No prod write by default Sandbox SVID has zero write capability to release branches / prod; writes only via Policy Server + signed release role + human control Egress control Per-sandbox egress allow-list (§2); browser/term tools route through the inspecting proxy Terminal & browser isolation term and browser.fetch tools run in a separate isolation domain; fetched content is untrusted input subject to context hygiene (§5) and cannot self-execute Secrets via Vault No long-lived secrets in env/image; HashiCorp Vault brokers short-TTL, narrowly-scoped, dynamic credentials; Vault audit log feeds the WORM store Kill-switches Per-agent and fleet-wide kill-switch: revoke SVID (SPIRE) → mesh denies all calls instantly; circuit-breakers on anomalous tool-call rate; "freeze on novel egress" tripwire Resource bounds CPU/GPU/wall-clock/tool-call quotas — caps runaway loops (cost + blast radius; see 08 ) 5. Prompt-injection, context hygiene & PII/PHI protection # Treat all model-facing content not authored by the harness as hostile input. Input sanitization & provenance. Retrieved context is wrapped with provenance and trust labels; instructions embedded in data (issues, comments, fetched docs) are demarcated and not treated as commands. Retrieval is restricted to a source allow-list (approved internal repos/doc stores); arbitrary web/repos are off the path for regulated work, eliminating most context-poisoning vectors. Placeholder injection — the [[VAR]] pattern (context hygiene middleware). Before any PHI/PII/secret-bearing content enters a prompt, a deterministic middleware masks sensitive spans into typed placeholders and keeps the mapping in a secure side-table the model never sees: text RAW: Patient John Doe (MRN 55512) reports error E13 at 10.0.4.7 MASKED: Patient [[NAME_1]] (MRN [[ID_1]]) reports error E13 at [[IP_1]] side-table (Vault-sealed): NAME_1→"John Doe", ID_1→"55512", IP_1→"10.0.4.7" The model reasons over placeholders; on output, only authorized placeholders are re-hydrated, and only into allow-listed sinks . PHI never reaches the model, never lands in logs/eval sets in cleartext, and cannot leak through the model channel. Output sanitization & insecure-output-handling defense. Agent output destined for a downstream interpreter (shell, SQL, code) is schema-validated and never auto-executed without passing the Policy Server; outputs are scanned for residual PII and for re-injection patterns. Defense against poisoned repos/docs. Source allow-listing + signed dependencies (§6) + semantic injection detection (§3.2) + the rule that data is never instruction . A poisoned README cannot redirect the agent's authority because authority lives in structural policy, not in text. The "rogue agent emails 50 colleagues" failure class. Worked example of defense-in-depth: Structural deny: email.send is not in any dev-agent role (§3.1) — the tool literally cannot be invoked. Even if a privileged role had it: semantic gate flags bulk-recipient/exfil intent → escalate. Egress allow-list: the SMTP endpoint is not reachable from the sandbox. Kill-switch: anomalous tool-call burst trips the circuit-breaker and revokes the SVID. Evidence: the attempt is recorded as a security incident (§11). No single control is trusted; the action requires all of them to fail simultaneously. 6. Supply-chain assurance # Provenance is required for code, models, AND datasets — models and data are first-class regulated supply-chain artifacts. Aligns to FD&C Act §524B and FDA premarket cybersecurity. Artifact Signing Provenance SBOM Admission check Code / container images Sigstore/cosign SLSA build provenance (L3 target) CycloneDX SBOM Gatekeeper verifies signature + provenance Model weights + adapters cosign-signed digest Build/train provenance (base model lineage, fine-tune run ID) Model SBOM (base model, datasets, hyperparams, eval hash) Unsigned/unknown model rejected at admission Datasets cosign-signed manifest + hash Source lineage, consent/PHI-handling attestation Dataset card / data SBOM Untrusted dataset cannot enter a training run Eval gold sets signed, version-pinned provenance to authoring QA included tamper = gate integrity incident ( 05 ) Admission enforcement. OPA/Gatekeeper admission policy: no pod runs a container or loads a model whose cosign signature and SLSA provenance do not verify against the trusted key set (Vault/SPIRE-rooted). Reproducible builds (P7) mean any shipped artifact — code or model — can be regenerated bit-for-bit and defended in an audit or recall. Vulnerability management (continuous SBOM scanning, KEV/CVE feeds) satisfies the §524B postmarket obligation for Track B artifacts and the QS obligation for Track A tooling. 7. Records, audit & 21 CFR Part 11 # Per P4 , every agent action is evidence . The evidence record is the regulatory product of the agent, not a byproduct. What is recorded for every step (immutable, attributable, replayable): Field Source Part 11 role Prompt + full context bundle (hashed; PHI masked) harness reconstructs what the agent saw Model tier + weights digest + adapter version registry (§6) "which software produced this" Tool call + args (sanitized) + Policy Server verdict Policy Server (§3) authorization record Verifier/eval results gates ( 05 ) objective evidence of correctness Human decision + e-signature (who, when, meaning) review system 21 CFR Part 11 §11.50/11.70 SPIFFE identity of every actor SPIRE attribution / non-repudiation Policy version + semantic-model version Policy Server change-control linkage Storage: WORM / immutable store, hash-chained (append-only, tamper-evident), time-synced. Replayability: because weights, adapters, prompts, and policy are all versioned and signed, any decision can be deterministically re-derived for an investigator. e-signatures bind a human's identity, timestamp, and the meaning of their action (reviewed / approved / authorized) to the record. Audit & recall use: in a recall or FDA inspection, the WORM store answers "show me everything the agent did to this Class C module, who authorized it, what it saw, and prove it wasn't tampered" — with cryptographic non-repudiation. This is the evidentiary backbone the CSA validation (§8) certifies. 8. Validating the agent as regulated software (CSA) # Under FDA Computer Software Assurance , the agentic harness is production/Quality-System software and is validated risk-proportionately — not exhaustively, but where it matters. CSA element Application here Intended use Defined per agent role (e.g., "propose unit tests for Class A modules"); autonomy bounded by §10 matrix Risk-based assurance Test effort scales with the impact of the agent's failure; Class C-touching agents get the deepest scrutiny (P3) Security testing Threat-led: each §1.3 threat has corresponding adversarial tests and red-team coverage (§11) Threat-led validation Validation cases derived from the threat model + OWASP LLM / ATLAS mappings, not just happy-path Objective evidence The §7 WORM record + the 05 eval evidence constitute validation evidence Change control Harness, policies, and model versions are controlled items; ISO/IEC 42001 governs the AI management system The assurance argument ties directly to 05-evaluation-and-validation : deterministic eval + ≥99.9% release-gate correctness is the functional assurance; this document supplies the security assurance. Together they form the CSA validation package. 9. Two regulated tracks # This is the distinction most frequently muddled — and the one an auditor will test. Dimension Track A — AI that BUILDS the device (this framework) Track B — AI shipped INSIDE the device (SaMD / AI function) What it is Dev/test/doc agents = production & QS tooling The model is part of the medical device / its output is a device function Submission-bearing? No — not in the 510(k)/PMA submission as a function Yes — part of premarket submission Primary regime CSA , ISO 13485/QMSR, 21 CFR Part 11 IEC 62304, ISO 14971, FDA premarket cyber, EU AI Act high-risk , PCCP Change control QS change control; ISO/IEC 42001 Predetermined Change Control Plan (PCCP) — pre-authorized model-update envelope Clinical evidence Not required Required (clinical validation of the AI function) Failure consequence Bad tooling → defective product (caught by gates) Bad model → direct patient harm in the field Shared assurance muscles (build once, apply to both): self-hosted signed model supply chain (§6), immutable evidence + Part 11 records (§7), threat-led validation (§8), drift/anomaly monitoring (§11), ISO/IEC 42001 AI governance. Where obligations diverge: Track B additionally owns clinical validation, a PCCP, premarket cybersecurity documentation, and EU AI Act high-risk conformity. This document governs Track A; it deliberately reuses controls that a Track B program will also need, but Track B's submission obligations are out of scope here. 10. Autonomy Authorization Matrix (canonical) # This is the canonical autonomy matrix. It is referenced by 02-maturity-model and 06-agentic-workflows . It maps (ASMM-Med governing level × IEC 62304 safety class) → permitted agent action and required human control. Per P3, Class C is ALWAYS dual human control regardless of maturity level. Action legend: Suggest (advisory only) · Propose-PR (opens a PR, no merge authority) · Auto-bounded (autonomous within signed, pre-authorized bounds) · Forbidden . Human control legend: None · Single review · Dual control (two qualified humans; author ≠ approver). ASMM-Med level ↓ / IEC 62304 class → Class A (no injury) Class B (non-serious injury) Class C (death / serious injury) L0 Ad-hoc Suggest / None Suggest / Single review Suggest / Dual control L1 Governed Assistance Suggest / None Propose-PR / Single review Propose-PR / Dual control L2 Spec-Driven Bounded Propose-PR / Single review Propose-PR / Single review Propose-PR / Dual control L3 Orchestrated Agentic Auto-bounded / Single review (post-hoc) Propose-PR / Single review Propose-PR / Dual control L4 Validated Autonomous Auto-bounded / None within bounds Auto-bounded / Single review Propose-PR / Dual control L5 Self-Optimizing Auto-bounded / None within bounds; sampled audit Auto-bounded / Single review Propose-PR / Dual control Reading the matrix: The leash lengthens with maturity (rows) but is capped by safety class (columns). Class C never reaches Auto-bounded or "None." The highest Class C autonomy is Propose-PR under dual control — the agent proposes and evidences; two qualified humans author the merge decision (P3). "Auto-bounded" requires the bounds to be signed, version-controlled policy enforced by the Policy Server (§3); outside the bounds, the agent escalates. Every cell's enforcement is mechanical: the Policy Server reads (role→level, target→safety_class) and applies the corresponding require: control (§3.1). 11. Continuous security # Security is a steady-state operation, not a one-time gate. Capability Implementation Red-team agents Standing adversarial agents continuously attempt prompt injection, context poisoning, tool misuse, and exfil against the live harness; findings feed §8 validation and §3 policy Adversarial eval OWASP-LLM / ATLAS-derived adversarial suites run in the deterministic eval pipeline ( 05 ); regressions block model/harness promotion Drift & anomaly response Monitor tool-call distributions, egress patterns, semantic-gate hit rates, and model-output drift; anomalies trip circuit-breakers (§4) and open incidents Incident handling Defined runbooks: SVID revocation, namespace freeze, fleet kill-switch, WORM-log forensic replay; incidents link to QMS CAPA Secure model-update path New weights/adapters: signed → SBOM'd → SLSA-provenanced (§6) → adversarial + functional eval gates ( 05 ) → Gatekeeper admission → staged rollout with rollback. For Track B models, this path executes within the PCCP envelope ; for Track A , under QS change control + ISO/IEC 42001 Vulnerability management Continuous SBOM/CVE scanning of code + model dependencies; §524B-aligned triage and disclosure Appendix A — Control-to-standard traceability # Control (this doc) Standard / framework anchor Policy Server, autonomy matrix (§3, §10) IEC 62304 §5–§9, P3; CSA Evidence / WORM records, e-signature (§7) 21 CFR Part 11 , ISO 13485/QMSR Zero-trust, mTLS, egress, sandboxing (§2, §4) IEC 62443 , FDA premarket cybersecurity Supply chain: signing/SBOM/SLSA (§6) FD&C Act §524B , SLSA, Sigstore Threat model, red-team, adversarial eval (§1, §11) OWASP LLM Top 10 , MITRE ATLAS AI management system, change control (§8, §11) ISO/IEC 42001 PII/PHI masking, context hygiene (§5) HIPAA, GDPR, ISO 14971 (risk) Track A vs B, PCCP, high-risk (§9) EU AI Act , EU MDR, FDA PCCP guidance Cross-references: autonomy bounds and maturity levels — 02-maturity-model ; harness/sandbox architecture — 03-reference-architecture ; model/adapter signing and fleet — 04-model-strategy-and-finetuning ; deterministic eval and validation evidence — 05-evaluation-and-validation ; workflow-level human controls — 06-agentic-workflows ; cost of controls — 08-token-and-gpu-economics . ← Previous 06 · Agentic Workflows Next → 08 · Token & GPU Economics On this page 0. Purpose and framing 1. Threat model for agentic dev in a regulated org 2. Zero-trust architecture 3. The Policy Server in depth 4. Agent sandboxing & blast-radius control 5. Prompt-injection, context hygiene & PII/PHI protection 6. Supply-chain assurance 7. Records, audit & 21 CFR Part 11 8. Validating the agent as regulated software (CSA) 9. Two regulated tracks 10. Autonomy Authorization Matrix (canonical) 11. Continuous security Appendix A — Control-to-standard traceability
# 08 · Token & GPU Economics · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/08-token-and-gpu-economics.html
## 08 — Token & GPU Economics (FinOps for Self-Hosted Agentic Dev) #
08 · Token & GPU Economics · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate 08 — Token & GPU Economics (FinOps for Self-Hosted Agentic Dev) # Part of Agentic-Native SDLC for Regulated Medical Device Engineering. Status: Reference baseline · Date: May 2026 · Audience: CFO/Finance, Platform Engineering, Quality/Regulatory. Cross-refs: 01-requirements · 02-maturity-model · 03-reference-architecture · 04-model-strategy-and-finetuning · 05-evaluation-and-validation · 06-agentic-workflows · 07-security-and-compliance · 09-adoption-roadmap This document is the financial control plane for the program. Every formula, rate, and ratio below is illustrative ; the parameters are org-set and owned by the FinOps practice. Where a number appears, treat it as a placeholder to be replaced by measured values from your own fleet telemetry. 1. The Economic Reframing # Self-hosting fine-tuned open-weight models is not primarily a performance decision for this program — it is an economic and sovereignty decision, and it changes the shape of the cost, not just the magnitude. A SaaS LLM API is pure, uncapped OpEx : a per-token meter that scales linearly and forever with usage, with no asset on the balance sheet and no floor on marginal cost. A 1000+ developer org running agentic workflows generates enormous token volume — most of it on repair loops, retrieval, and evaluation, not the final answer the human sees. At that volume, the per-token meter becomes the dominant line item and is structurally unbounded. Self-hosting converts that into two different categories: Dimension SaaS API (rejected) Self-Hosted Fleet (chosen) Cost class OpEx, per-token, uncapped GPU CapEx (amortized) + Ops OpEx Marginal cost of one more token Vendor rate (fixed, never zero) ≈ marginal electricity once GPU is owned Cost behavior at scale Linear, unbounded Step-function (buy capacity) + high utilization wins Data path Source/spec leaves the boundary Stays inside the regulated boundary Reproducibility Vendor-controlled model drift Pinned weights, P7 reproducible Negotiating position Vendor pricing power Internal control of the curve Why the org chose this (decision is non-optional, see §6): Cost at scale. Above a break-even volume, owned-and-amortized GPU beats per-token billing, and our volume is far above it. Data sovereignty. Regulated source, specifications, defect data, and patient-adjacent context must not transit a third-party LLM (see 07-security-and-compliance ). No per-token vendor billing / no model drift. Validation under IEC 62304 requires that a release gate run today reproduces tomorrow; a silently-updated vendor model breaks that (P7, 05-evaluation-and-validation ). The governing metric: COST-PER-GREEN-PR # Per Principle P6 , cost is measured per verified task , not per token. The unit of economic value this program produces is a Green PR : a change that passes all deterministic gates plus human review (see 06-agentic-workflows ). Tokens that do not contribute to a Green PR are waste, regardless of how cheap each one was. COST-PER-GREEN-PR = Total compute $ (inference + eval + retrieval + idle + amortized CapEx + ops) Number of PRs that passed all gates + review Why optimizing cost-per-token alone is a trap. Cost-per-token is a component , not the objective. The classic failure: a team cuts token price 40% by quantizing aggressively or routing everything to a tiny model, first-pass yield collapses, agents loop and re-attempt, escape-rate to human reviewers rises — and cost-per-Green-PR goes UP even as cost-per-token went down. Cheap-but-wrong is the most expensive mode in a regulated SDLC, because rework, re-validation, and reviewer time dwarf raw inference. The denominator is the lever; the numerator is the temptation. 2. Cost Taxonomy # All program compute spend decomposes into seven drivers. Each is metered separately (OpenTelemetry → cost, §7) so it can be attributed and optimized independently. # Cost driver What moves it Primary controls 1 Inference — model size Params served, tier (S/M/L), MoE active params Smallest-capable-model, tiered routing (§4), distillation 2 Inference — context length Prompt tokens, retrieved context, history Context economy (§4), prefix/prompt cache, retrieval over stuffing 3 Inference — output length Generated tokens, reasoning-effort, loop count Reasoning-effort caps, loop-count limits, structured output 4 Inference — batch efficiency Continuous batching, concurrency, queue depth vLLM continuous batching, PagedAttention, request shaping 5 Inference — GPU type/utilization GPU SKU $/hr, MIG slicing, idle fraction Right-sizing, MIG partitioning, KEDA scale-to-zero 6 Training / fine-tuning LoRA vs full FT, dataset size, epochs, runs Multi-LoRA, spot/preemptible + Kueue, distillation runs 7 Evaluation / validation Gate suites, reproducibility reruns, 99.9% sampling Eval caching, deterministic seeds, eval-tier routing — Retrieval / embedding Embedding calls, rerank, index refresh, vector store Tier-E batching, retrieval cache, incremental indexing — Idle / overprovisioning Reserved-but-unused GPU, warm pools, headroom Scale-to-zero, autoscaling, batch/interactive split — Ops / people Platform SRE, FinOps, MLOps, on-call, eval engineering Automation, self-service, maturity (L1→L5) Do not forget evaluation cost. The ≥99.9% release-gate correctness target (P1) is enforced by running a great deal of inference — large eval suites, adversarial probes, statistical sampling, and reproducibility reruns that re-execute gates on pinned weights for the regulatory record. For mature agentic repos, eval + reproducibility compute is frequently 25–45% of total inference spend and must be a first-class budget line, not an afterthought (see 05-evaluation-and-validation ). 3. An Illustrative Cost Model # ILLUSTRATIVE — all rates are placeholders. Parameters ( $/GPU-hr , throughput, yields) are org-set and replaced by measured fleet telemetry. The structure is the deliverable, not the digits. 3.1 GPU-hour → cost per 1k tokens, by tier # The per-token cost of a served model is the GPU rental cost divided by how many tokens that GPU produces per hour. Cost_per_1k_tokens = (GPU_count × $/GPU-hr) ÷ Utilization Effective_throughput_tok_per_hr × 1000 Effective_throughput already bakes in quantization, continuous batching, and speculative decoding (§4). Illustrative steady-state rates: Tier Model class GPU footprint (illus.) Eff. throughput (tok/s, batched) $/GPU-hr (illus.) $ / 1k tok (illus.) Tier-S "Reflex" 1–8B, FP8/INT8 MIG slice / 1× GPU 6,000 $2.50 $0.00012 Tier-M "Worker" 14–34B, AWQ/GPTQ 1–2× GPU 2,200 $2.50 $0.00063 Tier-L "Reasoner" 70B+/MoE 4–8× GPU 700 $2.50 $0.0079 Tier-V Multimodal VLM 1–2× GPU 1,000 $2.50 $0.0028 Tier-E Embed/Rerank embedding MIG slice 40,000 (items) $2.50 $0.0000175 The ~ 65× spread between Tier-S and Tier-L is the entire economic argument for tiered routing: a call needlessly sent to Tier-L costs as much as ~65 correct Tier-S calls. 3.2 Cost per agent task # A single agent task is rarely one model call. It is a sequence of calls across tiers, plus the verifier/sandbox compute that makes the work verifiable , plus its share of evaluation — all divided by first-pass yield (FPY) to account for repair loops. Cost_per_task = Σ_over_tiers( tokens_tier × rate_tier ) + Verifier_sandbox_compute + Eval_amortized First_Pass_Yield (0 < FPY ≤ 1) Verifier_sandbox_compute = build/test/static-analysis/sandbox-exec cost to check the candidate (the harness is the product, P5). Eval_amortized = task's share of gate + reproducibility runs. FPY is the multiplier that ties quality to cost : every failed attempt re-spends the numerator. The repair-loop explosion, holding raw token cost constant, as FPY falls: First-Pass Yield Effective cost multiplier (1 ÷ FPY) Interpretation 0.90 1.11× Healthy; small rework tax 0.70 1.43× Noticeable loop spend 0.50 2.00× Half of all work is redone 0.30 3.33× Loop-dominated; cheap model is a false economy 0.15 6.67× Pathological; escape-rate to humans spikes This table is the quantified form of the §1 trap: driving down per-token cost while letting FPY fall is a net loss. 3.3 Worked numeric example (labeled placeholders) # ILLUSTRATIVE. "Implement a bounded requirement-to-code change with passing unit tests." Input (placeholder) Symbol Value Tier-S router/classify + lint tokens t_S 8,000 tok @ $0.00012/1k Tier-M implementation tokens t_M 40,000 tok @ $0.00063/1k Tier-L escalation (10% of tasks need it) t_L 6,000 tok @ $0.0079/1k × 0.10 Embedding/retrieval t_E 20,000 items @ $0.0000175/1k Verifier/sandbox compute (build+test+sast) C_v $0.018 Eval/repro amortized share C_e $0.012 First-pass yield FPY 0.70 Token + retrieval cost: Tier-S : 8,000/1000 × $0.00012 = $0.00096 Tier-M : 40,000/1000 × $0.00063 = $0.02520 Tier-L : 6,000/1000 × $0.0079 × 0.10 = $0.00474 Tier-E : 20,000/1000 × $0.0000175 = $0.00035 Subtotal tokens = $0.03125 Numerator = $0.03125 + C_v($0.018) + C_e($0.012) = $0.06125 Cost_per_task = $0.06125 ÷ FPY(0.70) = $0.0875 ✅ Sensitivity — same task, FPY collapses to 0.30 (e.g., over-aggressive quantization or routing too small): $0.06125 ÷ 0.30 = $0.2042 — a 2.3× cost increase with zero change to per-token rates. If that low FPY also raises human escape-rate, the true cost-per-Green-PR rises further still (reviewer minutes are the most expensive tokens in the system). 4. The Optimization Levers # Each lever lists mechanism → expected impact (illustrative) → tradeoff . They compound; they also interact (over-using one can sink FPY and undo another), so they are tuned against cost-per-Green-PR, never in isolation. 4.1 Tiered model routing # Mechanism. A lightweight classifier/router running on Tier-S scores incoming task complexity and dispatches to the smallest capable tier; escalate to Tier-M/Tier-L only on confidence/complexity thresholds or verifier failure. Smallest-capable-model principle. Impact. If the majority of low-complexity calls resolve on Tier-S (~65× cheaper than Tier-L), blended $/token can drop 40–70% vs. always-on Tier-L. Tradeoff. Router error is double-edged: under-routing tanks FPY (loops); the router itself must be validated and is an eval surface. Mis-tuned thresholds look cheap per-token while raising cost-per-Green-PR. 4.2 Caching (KV/prefix, prompt, semantic, retrieval) # Mechanism. PagedAttention/KV-cache reuse + prefix/prompt caching skip recompute of shared system prompts, specs, and skill preambles; semantic caching returns prior answers for near-duplicate requests; retrieval caching avoids re-embedding/re-fetching stable context. Impact. Prompt/prefix cache can cut prefill compute 30–80% on repetitive agentic prompts (large shared spec/skill prefixes); retrieval cache cuts Tier-E load materially. Tradeoff. Semantic cache must be conservative in regulated paths — a stale or near-miss hit that flips a gate decision is a correctness defect. Cache keys must include weight/version/spec hashes for reproducibility. 4.3 Quantization + speculative decoding + continuous batching # Mechanism. FP8/INT8/AWQ/GPTQ shrink memory/raise throughput; speculative decoding uses a small draft model to propose tokens a larger model verifies; continuous batching (vLLM) keeps GPUs saturated across concurrent requests. Impact. Quantization commonly yields 1.5–3× throughput/$; speculative decoding 1.5–2.5× latency/throughput on accept-heavy workloads; continuous batching lifts utilization from ~30% to 70–90% . Tradeoff. Quantization can degrade accuracy — every quantized model must re-pass deterministic gates ( 05 ) before serving. Quantize, then measure FPY , never assume. 4.4 Multi-LoRA adapter amortization # Mechanism. Serve one base model with many LoRA adapters (per-domain/per-task behaviors) hot-swapped per request, instead of standing up a full fine-tuned model per behavior. Impact. Collapses N dedicated deployments into ~1 base footprint — large reduction in idle/overprovisioning and CapEx; new specialized behaviors become near-zero marginal serving cost. Tradeoff. Adapter routing/versioning complexity; per-adapter eval still required; a bad base upgrade invalidates all adapters at once (manage via 04 ). 4.5 Context economy # Mechanism. Lean specs; dynamic context / skills loaded on demand; retrieval over context-stuffing; prune history; structured rather than verbose I/O. Avoid "context dumping" entire repos/specs into every prompt. Impact. Context length drives prefill cost super-linearly via attention; trimming 50% of tokens often cuts prefill cost >50% and improves FPY (less distraction). Tradeoff. Under-supplying context tanks FPY too — economy means right context, not minimal context. Tune against yield. 4.6 Autoscaling, partitioning, queueing # Mechanism. KEDA scale-to-zero for spiky/interactive services; MIG partitioning to pack small models onto GPU slices; spot/preemptible for training/eval batch; Kueue for queue + quota fairness. Impact. Scale-to-zero eliminates overnight idle on bursty endpoints; MIG raises packing density; spot cuts training $ 60–90% . Tradeoff. Cold-start latency on scale-from-zero (mitigate with warm minimums for interactive tiers); spot preemption requires checkpointing. Never put latency-critical interactive gates on pure scale-to-zero without a warm floor. 4.7 Distillation (big → small) # Mechanism. Distill Tier-L behavior into Tier-S/M adapters; the expensive Reasoner generates training signal once , the cheap model serves it forever . Impact. Shifts steady-state load down a tier — recurring 40–65% serving-cost reduction on distilled task families; reduces Tier-L invocation frequency. Tradeoff. Up-front distillation + eval CapEx; distilled model can lag base capability on edge cases — gated re-validation required before it replaces escalation paths. 4.8 In-loop budget guardrails # Mechanism. Hard token/compute caps per task , reasoning-effort caps , and loop-count limits enforced by the orchestrator ( 06 ); on breach, fail-closed to human triage rather than burning unbounded compute. Impact. Bounds the worst-case tail — caps the cost of pathological low-FPY tasks that would otherwise loop indefinitely; protects the monthly budget from a single runaway agent. Tradeoff. Caps set too tight truncate legitimately hard tasks (raising escape-rate); caps are themselves tuned against cost-per-Green-PR. flowchart TD A[Incoming agent task] --> B{Cache hit?
prompt / semantic / retrieval} B -- yes --> Z[Return cached / cheap path] B -- no --> C[Tier-S router/classifier] C -->|low complexity| D[Tier-S Reflex
+ multi-LoRA adapter] C -->|medium| E[Tier-M Worker] C -->|high / escalated| F[Tier-L Reasoner
sparingly] D --> V{Verifier / deterministic gates} E --> V F --> V V -- pass --> G[GREEN PR candidate] V -- fail --> H{Budget guardrail:
tokens / loops / effort left?} H -- yes --> C H -- no --> I[Fail-closed → human triage] G --> M[OpenTelemetry cost metering →
cost-per-Green-PR] 5. Capacity Planning for 1000+ Developers # Sizing the fleet is a queueing problem, not a headcount multiplication. The goal is enough capacity to hold interactive latency SLOs at peak while keeping steady-state utilization high. Estimating concurrent load (illustrative). Active_devs = 1000 × engagement_factor(0.6) = 600 Req_per_active_dev_hr = 30 (agentic calls incl. loops/retrieval/eval) Average_RPS = 600 × 30 / 3600 ≈ 5 RPS sustained Peak_RPS = Average_RPS × peakiness(3.0) ≈ 15 RPS GPU_needed_at_peak = Peak_RPS ÷ per-GPU_throughput_at_SLO (per tier) Planning dimension Approach Peak vs. average Size interactive tiers for peak RPS at the latency SLO; size batch (eval/training) for average throughput with queueing. Peakiness factor measured per region/timezone. GPU fleet sizing Per-tier: ceil(Peak_RPS ÷ throughput_at_SLO) + headroom. Bottom-heavy fleet (mostly Tier-S/M, few Tier-L) mirrors the routing distribution. Batch vs. interactive separation Dedicated pools. Interactive = warm, latency-bounded, KEDA with warm floor. Batch = Kueue-queued, spot-backed, scale-to-zero, latency-tolerant. Never let a training job preempt an interactive gate. Multi-tenancy fairness Kueue quotas per team/repo so no tenant starves others; borrowing from idle quotas allowed, reclaimable on demand. Headroom for eval/training Reserve explicit capacity (illus. 15–25% ) for gate suites, reproducibility reruns , and fine-tuning — these are non-optional regulatory load, not discretionary. 6. Build-vs-Buy / Self-Host Math # ILLUSTRATIVE structured comparison. Numbers are placeholders to frame the reasoning , not a quote. Factor Hypothetical SaaS API at this scale Self-Hosted Fleet Annual token volume (illus.) ~30B billable tok/yr (incl. loops, eval, retrieval) same workload, owned compute Unit basis blended vendor $/1k tok amortized $/GPU-hr + ops Annual run cost (illus.) 30M × $X_blended_per_1k → large, uncapped, linear GPU_CapEx ÷ amort_yrs + Ops_OpEx + power Marginal next-token cost vendor rate (never zero) ≈ marginal power on owned GPU Cost trajectory at growth scales with usage forever flattens as utilization rises Break-even reasoning. Self-host carries up-front CapEx (GPUs, networking) + steady Ops OpEx (SRE, FinOps, MLOps, power, eval engineering). API carries zero fixed cost but a per-token meter. There is a crossover volume above which amortized self-host is cheaper: Break-even when: (CapEx ÷ amort_years) + Ops_OpEx_annual + Power_annual < Annual_token_volume × Blended_API_rate_per_token → Self-host wins decisively once volume × API_rate exceeds fixed+ops cost. At 1000+ devs with loop/eval/retrieval amplification, our volume is FAR above crossover. Non-cost drivers that make self-host non-optional here (these hold even if the math were neutral): Sovereignty. Regulated source, specs, and defect data must not leave the boundary ( 07 ). IP control. Proprietary device engineering knowledge stays in-house; no third-party training on our data. Regulatory control. Pinned, reproducible weights for IEC 62304 validation; no vendor-driven model drift mid-gate (P7, 05 ). The economics make self-host attractive ; the regulatory and sovereignty constraints make it mandatory . 7. FinOps Operating Model # Ties to D7 in 02-maturity-model . FinOps is a standing practice, not a quarterly cleanup. Capability Implementation Cost attribution Every inference/eval/training call tagged with team, repo, agent, tier, adapter, task-id, cache-status via OpenTelemetry spans → cost pipeline. Attributable to cost-per-Green-PR per repo. Dashboards OTel → cost warehouse → dashboards: $/Green-PR, FPY, tier mix, cache hit-rate, GPU utilization, eval/repro share, idle %. Budgets & alerts Per-team monthly budgets; alerts at 70/90/100%; automatic in-loop guardrails (§4.8) enforce hard caps independent of dashboards. Showback / chargeback Showback by default (visibility, behavior change); chargeback for high-volume teams to internalize cost. Cost SLOs Explicit objectives, e.g. $/Green-PR ≤ target , eval share ≤ 40% , GPU util ≥ 70% , cache hit ≥ 50% , idle ≤ 10% . Breach triggers review. Quarterly optimization review Re-tune routing thresholds, quantization choices, cache policy, fleet mix, distillation candidates against measured FPY and $/Green-PR. Feeds the L5 closed loop (§8). FinOps register (illustrative) # ID Metric / control Illustrative target Owner Cadence FIN-01 Cost-per-Green-PR ≤ $T per repo class FinOps + Repo lead Weekly FIN-02 First-Pass Yield ≥ 0.70 Platform + Eval Weekly FIN-03 GPU utilization ≥ 70% Platform SRE Daily FIN-04 Idle / overprovision ≤ 10% Platform SRE Daily FIN-05 Cache hit-rate (prompt+semantic) ≥ 50% Platform Weekly FIN-06 Tier-L invocation share ≤ 10% of calls Routing owner Weekly FIN-07 Eval + reproducibility share ≤ 40% of inference $ Quality Monthly FIN-08 Spot usage on training ≥ 80% of training GPU-hr MLOps Monthly FIN-09 Budget breach incidents 0 unbounded-loop runaways FinOps Monthly 8. Cost Across the Maturity Levels # Unit economics improve as the org climbs ASMM-Med — not because tokens get cheaper, but because the denominator (Green PRs) grows and waste shrinks. Level Cost profile Unit economics & control $/Green-PR trend L0 Ad-hoc Untracked, sporadic; no fleet metering No attribution; cost-per-token invisible Unknown / uncontrolled L1 Governed Assistance Mostly interactive assist; basic metering begins Cost-per-token visible; FPY undefined; little caching High, noisy L2 Spec-Driven Bounded Automation Bounded tasks; routing + caching introduced $/Green-PR first measured; guardrails appear Declining, variable L3 Orchestrated Agentic Workflows Multi-step agents; eval load rises sharply Full tier routing, multi-LoRA, budget caps; eval share managed Stabilizing L4 Validated Autonomous Agents High-volume autonomous, heavy validation/repro Strong FPY; distillation steady-state; tight cost SLOs Low, predictable L5 Self-Optimizing Agentic Enterprise Closed-loop cost optimization System auto-tunes routing/quantization/cache/fleet against $/Green-PR; data-driven distillation pipeline Minimized, self-correcting L5 is the target end state: routing thresholds, quantization choices, cache policy, and distillation candidates are selected by the system from telemetry, continuously, against cost-per-Green-PR — with every change still passing deterministic gates ( 05 ). 9. Anti-Patterns # Anti-pattern Why it costs Counter Always-on big models Tier-L (~65× Tier-S) serving routine calls; idle Reasoner GPUs Tiered routing + smallest-capable-model + scale-to-zero (§4.1, §4.6) Context dumping Whole repos/specs into every prompt; super-linear prefill cost; lowers FPY Context economy, retrieval over stuffing, prompt cache (§4.5, §4.2) Unbounded agentic loops Pathological tasks loop forever, blow the budget on one runaway Hard token/loop/effort caps, fail-closed to human (§4.8) Optimizing tokens, ignoring escape-rate/rework Cheap-but-wrong: per-token down, FPY down, $/Green-PR up — the §1/§3 trap Govern by cost-per-Green-PR; measure FPY before/after every change Idle GPU sprawl Reserved-but-unused GPUs, warm pools nobody uses, no batch/interactive split Scale-to-zero, MIG packing, Kueue quotas, idle SLO (FIN-04) No caching Recompute identical prefixes/embeddings every call KV/prefix + prompt + semantic + retrieval cache (§4.2) Forgetting eval cost 99.9% gates + reproducibility runs un-budgeted; surprise overruns Treat eval/repro as first-class line (FIN-07), reserve headroom (§5) Quantize-and-pray Throughput up, accuracy silently down, gates start failing Re-pass deterministic gates after every quantization (§4.3) Bottom line. This program does not minimize cost-per-token; it minimizes cost-per-Green-PR while holding ≥99.9% gate correctness. Self-hosting gives us the cost curve and the sovereignty; the levers in §4, the capacity discipline in §5, and the FinOps loop in §7 keep that curve flat as the org scales L1→L5. Quality is the dominant cost lever — a higher first-pass yield is cheaper than any cheaper token. ← Previous 07 · Security & Compliance Next → 09 · Adoption Roadmap On this page 1. The Economic Reframing 2. Cost Taxonomy 3. An Illustrative Cost Model 4. The Optimization Levers 5. Capacity Planning for 1000+ Developers 6. Build-vs-Buy / Self-Host Math 7. FinOps Operating Model 8. Cost Across the Maturity Levels 9. Anti-Patterns
# 09 · Adoption Roadmap · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/09-adoption-roadmap.html
## 09 — Adoption Roadmap #
09 · Adoption Roadmap · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate 09 — Adoption Roadmap # Program: Agentic-Native SDLC for Regulated Medical Device Engineering Audience: Executive sponsors, PMO, Engineering / QA-RA / Security leadership, AI Governance Board Status: Planning baseline (May 2026). All dates expressed as relative quarters from program start (Q1 = first full quarter after charter approval). All thresholds, ranges, and timelines are planning placeholders subject to gate-review revision. Cross-references: 01-requirements.md , 02-maturity-model.md , 03-reference-architecture.md , 04-model-strategy-and-finetuning.md , 05-evaluation-and-validation.md , 06-agentic-workflows.md , 07-security-and-compliance.md , 08-token-and-gpu-economics.md . 1. Roadmap Philosophy # This roadmap is assurance-gated, not capability-gated . We do not deploy the most capable agent we can build; we deploy the most capable agent we can validate, secure, and govern — and no further. The maturity model (ASMM-Med, see 02 ) is the spine; this document is the sequencing, resourcing, and decision layer that moves the organization L0 → L5 without ever letting autonomy outrun control. Five non-negotiable design rules: # Rule Consequence for sequencing R1 Governing level = min(D1, D4, D6) Capability dimensions (D2, D3, D5) may sprint ahead in build, but the enabled autonomy level is clamped by the weakest of Governance (D1), Eval/Assurance (D4), Security (D6). R2 Don't grant autonomy you can't yet validate (P3 + P1) Each phase ships its eval/validation evidence before the corresponding agent privilege is unlocked in production. The harness leads; the agent follows. R3 Reversible by construction Every promotion has a defined rollback (feature-flag, model-version pin, privilege revocation, scope reduction). No one-way doors. R4 Pilot before scale Capabilities are proven on low-safety-class, low-blast-radius repos before any Class B/C or fleet-wide exposure. R5 Value per phase Every transition must unlock a defensible, measurable engineering or quality value — not just platform plumbing — or the phase is reconsidered. These rules operationalize the Seven Principles ( 02 ): P1 (99.9% release-gate correctness as a system property), P2 (determinism wraps probabilism), P3 (risk-proportional autonomy), P4 (Part 11 evidence), P5 (the harness is the product), P6 (cost-per-green-PR), P7 (self-hosted reproducibility). Promotion across any level requires tri-signature: Engineering + QA/RA + Security . Safety-class gating per IEC 62304 is absolute — Class C software changes are always under dual human control regardless of maturity level. 2. Phase Plan # The program is five sequenced transitions. Build work for dimension N+1 may begin while operating at level N , but the level is only declared after the exit gate passes. Illustrative trajectory (from 02 ): Q1–Q2 → L1; Q3–Q4 → L2; Q5–Q7 → L3; Q8–Q11 → L4; Q12+ → L5. gantt title Agentic-Native SDLC — ASMM-Med Adoption Roadmap (relative quarters) dateFormat YYYY-MM-DD axisFormat Q%q section L0→L1 Governed Assistance Kill shadow AI / policy baseline :a1, 2026-01-01, 60d Self-hosted serving + central logging :a2, 2026-01-15, 75d L1 exit gate (Eng+QA/RA+Sec) :milestone, m1, 2026-03-31, 0d section L1→L2 Spec-Driven Bounded Automation SDD + AGENTS.md + specs/ adoption :b1, 2026-04-01, 80d Deterministic eval harness in CI :b2, 2026-04-15, 90d L2 exit gate :milestone, m2, 2026-09-30, 0d section L2→L3 Orchestrated Agentic Workflows Sandboxed multi-step agents + MCP plane :c1, 2026-10-01, 110d Model fleet + routing + policy server :c2, 2026-10-15, 120d HITL workflow rollout :c3, 2026-12-01, 90d L3 exit gate :milestone, m3, 2027-06-30, 0d section L3→L4 Validated Autonomous Agents CSA agent validation + 62304 traceability:d1, 2027-07-01, 150d 99.9% gate hardening + A2A :d2, 2027-09-01, 160d L4 exit gate :milestone, m4, 2028-06-30, 0d section L4→L5 Self-Optimizing Enterprise Closed-loop fine-tuning + eval promotion :e1, 2028-07-01, 150d Cost-optimal routing + PCCP change ctrl :e2, 2028-09-01, 160d L5 steady state :milestone, m5, 2029-03-31, 0d 2.1 Phase A — L0 → L1: Governed Assistance (≈ Q1–Q2) # Objective. Eliminate ungoverned ("shadow") AI use; stand up self-hosted serving with complete prompt/response logging so that every AI interaction is observable, attributable, and policy-bound. This phase buys legitimacy , not autonomy. Workstreams by dimension. Dim Workstream D1 AI Use Policy ratified; AI Governance Board chartered; ISO/IEC 42001 AIMS scaffolding initiated; shadow-AI amnesty + cutover. D2 Stand up vLLM/Triton on K8s + GPU Operator; serve baseline open-weight S/M models; KServe endpoints; no fine-tuning yet. D3 Inventory knowledge sources; begin curated repo/context indexing; no agent retrieval yet. D4 Define eval taxonomy and baseline manual review checklist; no automated harness yet. D5 IDE assist only (completion/chat); no tools, no agency. D6 Egress controls to block external LLM SaaS; Vault for secrets; SPIFFE/SPIRE identity bootstrap; Istio mTLS baseline. D7 OpenTelemetry tracing of all inference; central prompt/response log (Part 11-style retention); first FinOps GPU dashboard. D8 Org-wide AI literacy training; name first Model Steward and Eval Owner; communicate the "structure scales, vibes don't" thesis. Deliverables / artifacts. AI Use Policy v1; serving platform runbook; inference audit log schema (Part 11-aligned); shadow-AI decommission report; GPU baseline capacity plan; Governance Board charter + RACI. Exit criteria (who signs off). 100% of sanctioned AI traffic routed through self-hosted endpoints; zero known external LLM SaaS egress (Security verified). Complete, immutable, attributable logging of all prompts/responses (QA/RA verified for Part 11 retention). Governance Board operational with documented decision cadence (Engineering + QA/RA + Security tri-sign). Governing level confirmed: min(D1,D4,D6) ≥ L1. Primary risks + mitigations. Shadow AI persists covertly → egress blocking + amnesty + monitoring + leader modeling. Serving instability erodes trust → conservative SLOs, gradual cutover, fallback. Logging seen as surveillance → no-blame framing, transparency on purpose (quality/regulatory, not individual performance). Value unlocked. A defensible, auditable AI footprint; the substrate (serving + logging + identity) on which everything else is validated. 2.2 Phase B — L1 → L2: Spec-Driven Bounded Automation (≈ Q3–Q4) # Objective. Make specification-driven development (SDD) and a deterministic eval harness in CI the default. Move from "AI helps me type" to "AI executes against a spec, and a deterministic gate judges the result." This is where the harness becomes the product (P5). Workstreams by dimension. Dim Workstream D1 Map SDD artifacts to design controls (DHF inputs); QMSR/ISO 13485 (Feb 2026) alignment of AI-assisted change records. D2 MLflow registry; first task-specific LoRA fine-tunes (S/M tier) on curated internal code/spec corpora; reproducible build pipeline. D3 AGENTS.md repo conventions; specs/ directory standard; curated retrieval over approved knowledge; context provenance. D4 Deterministic eval harness in CI (Argo / pipeline-triggered); golden datasets; pass/fail gates; seedable, reproducible scoring (P2, P7). D5 Bounded single-step automation: scaffold, test-gen, doc-gen — no multi-step autonomy, no tool plane yet. D6 Sigstore/cosign signing of model + harness artifacts; SLSA provenance + SBOM for the AI toolchain. D7 Cost-per-green-PR (P6) instrumented; eval run cost tracking; token budgets per pipeline. D8 First Harness Engineers chartered; "tests/evals before code" engineering norm; review-every-line discipline. Deliverables / artifacts. AGENTS.md + specs/ standard; deterministic eval harness (versioned, signed); golden eval datasets; LoRA fine-tune cards; SBOM + SLSA attestations; cost-per-green-PR baseline. Exit criteria (who signs off). Deterministic eval harness gates CI on pilot repos with reproducible (seed-stable) results across runs (Engineering + Eval Owner). SDD artifacts traceable into design-control records (QA/RA). All models/harness artifacts signed with provenance (Security). Governing level: min(D1,D4,D6) ≥ L2. Primary risks + mitigations. Eval flakiness/non-determinism → strict seeding, hermetic environments, quarantine of flaky cases. Spec quality varies → spec templates, peer review, "spec-as-design-input" training. Fine-tune overfit → held-out eval sets, eval-driven acceptance only. Value unlocked. Trustworthy automated quality gates; measurable cost-per-green-PR; the first hard evidence that AI output can be validated before merge. 2.3 Phase C — L2 → L3: Orchestrated Agentic Workflows (≈ Q5–Q7) # Objective. Introduce sandboxed multi-step agents with a governed MCP tool plane , a model fleet with routing , a policy server , and human-in-the-loop (HITL) controls. Agents now plan and act across steps — inside sandboxes, under policy, with a human approving consequential actions. Workstreams by dimension. Dim Workstream D1 Policy-as-code (OPA/Gatekeeper + policy server) encodes who/what/where agents may act; risk-proportional autonomy matrix by safety class. D2 Model fleet tiers S/M/L/V/E operational; Ray/Kueue scheduling; multi-LoRA serving; routing by task/cost/quality. D3 Agent-grade retrieval + MCP-exposed knowledge resources; context windows scoped per task and per safety class. D4 Eval extended to trajectory and tool-use evaluation; HITL decision logging feeds eval; assurance cases per workflow. D5 Argo Workflows orchestration; MCP tool plane; gVisor/Kata sandboxing; Agent Stewards own each workflow; HITL checkpoints. D6 Zero-trust per agent identity (SPIFFE/SPIRE); least-privilege tool scopes; IEC 62443 alignment; egress-controlled sandboxes. D7 KEDA autoscaling on agent load; per-workflow FinOps; trajectory observability; token/GPU attribution per agent run. D8 Agent Steward + Harness Engineer roles scaled; HITL reviewer training; approval-fatigue controls designed (see §4). Deliverables / artifacts. MCP tool registry + scopes; policy server rulesets; agent sandbox runbook; per-workflow assurance case; routing policy; HITL approval logs; agent identity inventory. Exit criteria (who signs off). Multi-step agents run only in sandboxes with enforced least-privilege tool scopes (Security). Policy server denies out-of-scope actions by default; all consequential actions have HITL approval with audit trail (QA/RA + Security). Trajectory-level eval coverage meets threshold on pilot workflows (Eval Owner). Governing level: min(D1,D4,D6) ≥ L3. Class C remains dual-human. Primary risks + mitigations. Agent escapes sandbox / scope creep → default-deny policy, runtime sandbox, continuous policy tests. Tool-plane supply-chain risk → signed MCP servers, scoped credentials, Vault brokering. HITL becomes rubber-stamp → batched-but-meaningful approvals, sampling audits, no-blame escalation. Value unlocked. Real end-to-end task automation (multi-file changes, investigation, refactors) with human consequence-gating — the first order-of-magnitude productivity step, safely bounded. 2.4 Phase D — L3 → L4: Validated Autonomous Agents (≈ Q8–Q11) # Objective. CSA-validate agents as part of the QMS so that defined agent workflows can act autonomously (within safety class) at ≥99.9% release-gate correctness , with full IEC 62304 traceability and A2A (agent-to-agent) coordination. This is the regulated leap: agents become validated tools . Workstreams by dimension. Dim Workstream D1 CSA validation packages per agent; ISO 14971 risk analysis for agent failure modes; QMSR/13485 integration; §524B + cybersecurity documentation. D2 Locked, signed model+LoRA versions per validated workflow; reproducible serving; change control on model versions. D3 Validated knowledge sources; controlled context; provenance required for any retrieval feeding a Class B/C change. D4 ≥99.9% system-property gate demonstrated and continuously monitored; deterministic eval as validation evidence; assurance cases signed. D5 A2A coordination among validated agents; autonomy scoped strictly by safety class; Class C always dual human control. D6 Full IEC 62443 posture; cryptographic attestation of every agent action; tamper-evident audit. D7 Continuous gate-correctness monitoring; cost-per-green-PR optimized; drift + regression alarms. D8 QA/RA + Security embedded in agent lifecycle; Eval Owner owns validation evidence; operating model matured (§4). Deliverables / artifacts. Per-agent CSA validation report; IEC 62304 traceability matrix (requirement → design → agent action → test/eval → evidence); ISO 14971 agent FMEA; 99.9% gate-correctness monitoring dashboard; A2A protocol spec; signed assurance cases. Exit criteria (who signs off). Validated agents demonstrate ≥99.9% release-gate correctness as a sustained system property (Eval Owner + QA/RA). End-to-end IEC 62304 traceability for every autonomous action (QA/RA). CSA validation accepted into the QMS; reversibility + version pinning enforced (Engineering + QA/RA + Security). Class C dual-human control verified intact (Security + QA/RA). Governing level: min(D1,D4,D6) ≥ L4. Primary risks + mitigations. Regulator non-acceptance of agent validation approach → early FDA/CSA engagement, conservative assurance cases, pilot scope. 99.9% not met → no promotion; remain L3; harden harness. Drift erodes validated state → continuous monitoring + automatic rollback to pinned version. Value unlocked. Bounded autonomous engineering for lower-risk classes with regulatory-grade evidence — sustained throughput gains without sacrificing the audit trail. 2.5 Phase E — L4 → L5: Self-Optimizing Agentic Enterprise (≈ Q12+) # Objective. Close the loop: eval-driven, cost-optimal fine-tuning and promotion under PCCP-style change control . The system improves itself within pre-authorized bounds, with every change gated by the deterministic harness and governed change control. Workstreams by dimension. Dim Workstream D1 FDA AI/PCCP-style predetermined change-control protocol authored and approved; ISO/IEC 42001 AIMS at full maturity. D2 Closed-loop fine-tuning pipeline; candidate models auto-trained from production signal; promotion only via eval gate. D3 Self-curating knowledge with provenance + freshness controls; feedback-curated eval datasets. D4 Eval-driven promotion : a model/agent is promoted only if it beats incumbent on the deterministic harness at ≥99.9% (P1, P5). D5 Autonomous fleet self-optimization (routing, LoRA selection) within PCCP envelope. D6 Continuous attestation of self-modifying components; change provenance; rollback always available. D7 Cost-optimal routing (P6) closed-loop with FinOps; auto-rightsizing GPU; token economics steered to target. D8 Operating model steady-state; CoE → embedded; continuous enablement; no-blame, evidence-first culture institutionalized. Deliverables / artifacts. PCCP change-control protocol; closed-loop fine-tune pipeline; eval-driven promotion policy; cost-optimization control loop; AIMS conformance evidence. Exit criteria (steady state, who signs off). Every self-initiated model/agent change passes deterministic eval gate ≥99.9% before promotion, within PCCP envelope (Eval Owner + QA/RA). Cost-per-green-PR trending to target under FinOps governance (Engineering + PMO). All changes attested, reversible, and within pre-authorized change-control bounds (Security + QA/RA). Governing level: min(D1,D4,D6) ≥ L5. Primary risks + mitigations. Self-optimization drifts outside intended behavior → PCCP envelope as hard boundary; eval-gated promotion; rollback. Cost optimization degrades quality → quality is the gate, cost is the objective subject to the gate. Change control too slow → predetermined protocol pre-authorizes the space of changes. Value unlocked. A continuously improving, cost-optimal, self-hosted agentic SDLC where quality is provably non-decreasing and change is governed — the north star. 3. Pilot Strategy # Principle: prove on the safe edge, then graduate inward. Team / repo selection (in priority order): IEC 62304 Class A software first — non-safety internal tools, build tooling, test utilities, internal web apps. No patient-impact path. High test coverage + mature CI (the harness needs something to gate against). Volunteer teams with engaged tech leads (cultural readiness over raw size). Repos with clean, current specifications or willingness to write them. Explicitly excluded from early pilots: any Class B/C, regulated firmware, anything in a device's safety path. Success criteria for a pilot. Metric Target (placeholder) Deterministic eval gate reproducibility 100% seed-stable across reruns Defect-escape rate vs. baseline ≤ baseline (no regression) Cost-per-green-PR Measured + trending down Reviewer trust (survey) ≥ 70% "would expand scope" Rollback events causing incident 0 Blast-radius containment. Sandboxed execution (gVisor/Kata); least-privilege tool scopes; feature-flagged rollout; no production/clinical data; no write access to release branches without HITL; per-pilot kill switch (revoke agent identity via SPIFFE/SPIRE); model versions pinned and signed. Graduation path. Pilot → cohort (3–5 teams, same safety class) → broader Class A → cautious Class B only after the corresponding ASMM-Med level + eval evidence exist → Class C only with validated agents (L4) and always dual human control. Every graduation is a documented decision checkpoint (§8) with tri-signature. Learnings (harness components, specs, eval datasets, runbooks) are promoted to shared org assets owned by the CoE — the harness is the product (P5). 4. Organization & Operating Model Evolution # New / changed roles. Role Mandate Introduced Harness Engineer Builds/owns the deterministic eval harness, golden datasets, CI gates. Treats harness as a product. L2 Eval Owner Owns validation evidence, eval coverage, the 99.9% system property, promotion eval gates. L1 (named) → L2 (active) Model Steward Owns model fleet lifecycle, fine-tunes, versioning, signing, registry, reproducibility. L1 Agent Steward Owns a specific agent workflow: scope, policy, sandbox, HITL design, assurance case. L3 AI Governance Board Tri-functional (Eng + QA/RA + Security) authority over policy, promotions, gate reviews, stop/rollback. L1 QA/RA Integration Embeds regulatory/quality into the AI lifecycle: CSA validation, 62304 traceability, design controls. Throughout, deepening L2→L4 Security Integration Zero-trust agent identity, supply-chain, sandboxing, attestation, IEC 62443. Throughout RACI for promotion decisions (level N → N+1). Activity Eng Lead Eval Owner QA/RA Security Gov. Board PMO Produce eval/validation evidence C R C C I I Verify regulatory traceability I C R C I I Verify security posture I I C R I I Promotion decision (tri-sign) A C A A R C Stop / rollback trigger A C A A R I Resource / schedule C I I I C R/A (R=Responsible, A=Accountable, C=Consulted, I=Informed. Promotion requires the three A signatures: Eng + QA/RA + Security.) CoE vs. embedded. Start Center-of-Excellence (L1–L2): a small central team owns the harness, serving, policy, and standards. Transition to embedded (L3+): CoE retains shared assets, standards, and the Governance Board; Harness/Agent/Model Stewards embed in product teams. By L5, CoE is a thin standards-and-platform org; capability lives in teams. Scaling enablement to 1000+ devs. Train-the-trainer cohorts; AGENTS.md / specs/ as self-serve standards; golden-path templates; internal certification for HITL reviewers and Agent Stewards; office hours + internal community; documentation as code. Approval-fatigue & no-blame controls. Risk-proportional HITL (only consequential actions gated); batch low-risk approvals with audit sampling; clear escalation paths; rotation of reviewers; no-blame culture — logging is for quality/regulatory evidence, never individual performance; psychological safety to halt or roll back without penalty; "review every shipped line" framed as engineering craft, not blame. 5. Investment & Resourcing per Phase # Illustrative qualitative ranges (planning placeholders; cost mechanics per 08 ). Phase GPU capacity Platform/MLOps HC Fine-tuning effort Eval engineering Training/enablement L0→L1 Small (serving S/M, inference only) 3–6 None Manual/baseline Org-wide literacy (high reach, low depth) L1→L2 Small–Med (+ LoRA fine-tune jobs) 6–10 Moderate (task LoRAs) Heavy (harness is the product) Harness Engineer cohort; SDD training L2→L3 Med–Large (fleet S/M/L/V/E, routing) 10–18 Moderate–High High (trajectory/tool eval) Agent Steward + HITL reviewer training L3→L4 Large (validated serving + monitoring) 15–25 High (validated tunes) Very high (99.9% assurance + CSA) QA/RA + Security deep embed L4→L5 Large, cost-optimized (auto-rightsized) 12–20 (efficiency gains) Continuous (closed-loop) Continuous (eval-driven promotion) Steady-state continuous enablement Cost framing (ties to 08 ). GPU/token cost is a first-class constraint (P6). Early phases over-provision for trust; from L4→L5, FinOps + cost-optimal routing drive cost-per-green-PR down while quality (the gate) is held constant. Eval engineering is the largest sustained investment — the harness is the product, and validation evidence is the moat. Headcount shifts from central platform build (L1–L2) toward embedded stewardship + efficiency (L4–L5). 6. Consolidated Milestone & KPI Table per Phase # Capability KPIs and assurance/cost KPIs (drawn from 02 §8). Targets are placeholders. Phase Capability KPIs Assurance KPIs Cost KPIs Key milestone L0→L1 % AI traffic on self-hosted endpoints (→100%); AI literacy completion 100% prompt/response logged & attributable (Part 11) GPU baseline $/inference established Shadow AI killed; serving + logging live L1→L2 % pilot repos with SDD + AGENTS.md / specs/ ; automated change throughput Deterministic eval gate reproducibility (→100%); eval coverage % Cost-per-green-PR baseline Deterministic harness gates CI L2→L3 # sandboxed agent workflows; multi-step task completion rate Trajectory/tool-use eval coverage; HITL audit-trail completeness 100% $/agent-run; routing cost efficiency MCP plane + policy server + HITL live L3→L4 # validated autonomous workflows; autonomous PR throughput (by safety class) ≥99.9% gate correctness (system property) ; 100% IEC 62304 traceability Cost-per-green-PR optimized vs. L3 CSA-validated agents in QMS; A2A L4→L5 Closed-loop promotion frequency; fleet self-optimization rate Eval-driven promotion pass-rate ≥99.9%; PCCP-conformant changes 100% Cost-per-green-PR at target; GPU utilization Self-optimizing, PCCP-governed steady state 7. Program Risk Register # ID Risk Likelihood Impact Mitigation Owner PR-1 Capability outruns assurance (autonomy enabled before validation) Med Critical min(D1,D4,D6) clamp; no promotion without tri-sign + eval evidence; "don't grant autonomy you can't validate" AI Governance Board PR-2 Regulatory non-acceptance of agent validation / PCCP approach Med High Early FDA/CSA engagement; conservative assurance cases; CSA + IEC 62304 grounding; pilot scope QA/RA PR-3 Cost overrun (GPU/token) Med High FinOps from L1; cost-per-green-PR KPI; cost-optimal routing; rightsizing; tier S/M/L/V/E discipline PMO + Eng (FinOps) PR-4 Talent gap (Harness/Agent/Model Stewards, eval engineers) High High Train-the-trainer; CoE seeding; certification; phased role introduction; embedded model D8 lead / People PR-5 Shadow AI persists Med High Egress blocking; amnesty; monitoring; leader modeling; no-blame culture; make sanctioned path better Security PR-6 Model supply-chain compromise Low Critical Self-hosted open-weight only; Sigstore/cosign + SLSA + SBOM; signed LoRAs; Vault-brokered creds; attestation Security + Model Steward PR-7 Change-management resistance Med Med No-blame culture; value-per-phase wins; reviewer rotation; approval-fatigue controls; transparent comms Eng leadership PR-8 Eval non-determinism / flakiness Med High Hermetic envs; strict seeding; flaky-case quarantine; harness-as-product investment Harness Engineer / Eval Owner PR-9 Drift erodes validated state (post-L4) Med High Continuous gate-correctness monitoring; auto-rollback to pinned version; PCCP envelope Eval Owner 8. Decision Checkpoints & Governance Cadence # Cadence Forum Purpose Quarterly ASMM-Med Assessment Score all 8 dimensions; recompute governing level = min(D1,D4,D6); revise roadmap/thresholds Per transition Gate Review (tri-sign) Verify exit criteria; Eng + QA/RA + Security promotion sign-off; record reversibility plan Per pilot graduation Checkpoint Approve scope expansion / next cohort with evidence Continuous Monitoring + alarms 99.9% gate correctness, drift, cost, security posture Stop / rollback triggers (any one triggers halt + Governance Board review): Release-gate correctness drops below the level's threshold (e.g., <99.9% at L4). Any agent action outside policy/scope, or sandbox escape. Loss of attributable audit trail / Part 11 integrity. Cost-per-green-PR breaches FinOps ceiling without quality justification. Regulatory or QA/RA finding against a deployed capability. Model supply-chain or attestation failure. Rollback mechanics (always available, R3): feature-flag disable; pin to prior signed model/LoRA version; revoke agent identity (SPIFFE/SPIRE); reduce autonomy scope one ASMM-Med level; revert to HITL or dual-human control. No one-way doors. 9. "Start Monday" Quick Wins # Regulated adaptation of the source-paper spirit — structure, not vibes. For individual developers: Add an AGENTS.md to your repo: conventions, build/test commands, guardrails, what agents may and may not do. Create a specs/ directory; write the spec (the design input) before the code. Write tests and evals before code. The eval is the contract. Review every shipped line — AI-authored or not. Authorship is yours; craft is intent + validation. Route all AI use through sanctioned self-hosted endpoints. Kill your shadow AI today. For engineering leaders: Stand up (or adopt) the deterministic eval harness in CI for one repo this week. Treat the harness, golden datasets, specs, and runbooks as shared assets , not local hacks. Pick a Class A pilot repo with good coverage and a willing team. Model no-blame behavior: reward halting and rollback, not heroics. For the organization: Charter the AI Governance Board (Eng + QA/RA + Security). Name the first Eval Owner and Model Steward. Publish the AI Use Policy and the shadow-AI cutover plan. Begin org-wide AI literacy with the thesis up front: structure scales, vibes don't. 10. North-Star Vision Recap # The craft is changing, not disappearing. Intent and validation are the new engineering craft : a developer's value moves from typing implementation to specifying intent precisely and proving correctness rigorously. The harness is the product (P5); the eval is the contract; the 99.9% release gate is a system property , not a hope (P1). Structure scales; vibes don't. Spec-driven development, deterministic evaluation wrapping probabilistic generation (P2), risk-proportional autonomy (P3), and Part 11-grade evidence (P4) are what let a 1000+ engineer regulated organization adopt agentic SDLC without trading away the audit trail, the safety case, or patient trust. AI here is an amplifier of engineering and quality culture — never a substitute for it. Applied to a mature, evidence-first, no-blame culture, it compounds quality and throughput. Applied to a weak one, it compounds risk. This roadmap's discipline — assurance-gated, reversible, governed by min(D1,D4,D6), pilot-before-scale — is precisely how we ensure it amplifies the right thing. Don't grant autonomy you can't yet validate. Earn each level. Then the structure carries you to the next. ← Previous 08 · Token & GPU Economics Next → Assessment Scorecard On this page 1. Roadmap Philosophy 2. Phase Plan 3. Pilot Strategy 4. Organization & Operating Model Evolution 5. Investment & Resourcing per Phase 6. Consolidated Milestone & KPI Table per Phase 7. Program Risk Register 8. Decision Checkpoints & Governance Cadence 9. "Start Monday" Quick Wins 10. North-Star Vision Recap
# Assessment Scorecard · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/scorecard.html
## ASMM-Med Assessment Scorecard #
Assessment Scorecard · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate ASMM-Med Assessment Scorecard # A fillable, self-contained template for a quarterly maturity assessment. Select the satisfied level (0–5) for each of the eight dimensions and record the supporting evidence. Your governing level = min(D1, D4, D6) is computed live — capability can never outrun governance, evaluation, and security. Entries save to this browser; use Print to produce a PDF for the quality record. See the maturity model for full cell descriptors. Save Print / PDF Reset Governing level — Capability avg — Organization / Business unit Assessor(s) Assessment date D1 Governance, Quality & Regulatory governing QMS integration, IEC 62304 / ISO 13485 / QMSR alignment, CSA validation, traceability, change control L0 None / shadow AI L1 AUP + logging L2 QMS + SDD controlled L3 GAMP risk-class + 42001 L4 CSA-validated + traceable L5 PCCP + auto-evidence D2 Model Infrastructure & MLOps Self-hosted serving, fine-tuning pipeline, registry, reproducibility, multi-LoRA, GPU platform L0 None / SaaS L1 Single self-hosted L2 Registry + first fine-tunes L3 Tiered fleet + multi-LoRA L4 Reproducible validated pipelines L5 Continuous FT + auto-promote D3 Context & Knowledge Engineering Specs, rule files, RAG over code/docs/regulatory corpus, memory, context hygiene L0 Editor buffer only L1 System prompts L2 specs + AGENTS.md + RAG L3 Governed RAG + skills + hygiene L4 Validated sources + provenance L5 Self-curating + measured D4 Evaluation, Validation & Assurance governing Deterministic verifiers, eval suites, trajectory eval, the 99.9% gate, abstention L0 None / “looks right” L1 Lint + CI L2 Deterministic harness in CI L3 Output + trajectory evals L4 ≥ 99.9% gate + statistical L5 Continuous + closed-loop D5 Agentic Orchestration & Tooling Single → multi-agent, MCP/A2A, sandboxing, HITL design, workflow engine L0 None L1 Inline only L2 Single-step bounded L3 Multi-step sandboxed + MCP L4 Multi-agent A2A + hooks L5 Self-orchestrating D6 Security & Zero-Trust governing Identity, egress control, prompt-injection defense, supply chain, secrets, audit immutability L0 Uncontrolled L1 SSO + RBAC + DLP L2 Repo-scoped + Vault L3 Zero-trust + OPA + sandbox L4 Supply-chain + WORM + 62443 L5 Continuous red-team D7 Observability & FinOps Tracing, eval dashboards, token/GPU metering, routing economics, budget guardrails L0 None L1 Basic logs + GPU util L2 Usage + cost dashboards L3 Tracing + cost/task L4 Cost-per-green-PR SLO L5 Closed-loop cost opt D8 People, Skills & Operating Model Roles (conductor/orchestrator), review culture, training, approval-fatigue controls L0 Individual experiments L1 Training + champions L2 Spec/review skills L3 Orchestrator role emerges L4 Formal roles + failure-mode review L5 Judgment-first hiring Promotion rule. The overall operating level is the minimum across D1, D4, and D6. Advancing requires this scorecard plus documented evidence, signed by Engineering, Quality/Regulatory, and Security. Even at L4/L5, IEC 62304 Class C changes always require dual human control. Sign-off Engineering Quality & Regulatory Security ← Previous 09 · Adoption Roadmap Next → Diagrams (SVG)
# Diagrams (SVG) · Agentic-Native SDLC for Regulated MedTech
URL: https://unovie.ai/resources/medsdlc-html/diagrams.html
## Architecture & Workflow Diagrams #
Diagrams (SVG) · Agentic-Native SDLC for Regulated MedTech ← Unovie.AI ☰ ◆ Agentic-Native SDLC · Regulated MedTech Documents Overview Executive Brief 01 · Requirements 02 · Maturity Model 03 · Reference Architecture 04 · Model Strategy & Fine-Tuning 05 · Evaluation & Validation 06 · Agentic Workflows 07 · Security & Compliance 08 · Token & GPU Economics 09 · Adoption Roadmap Assessment Scorecard Diagrams (SVG) Self-hosted · FDA / IEC 62304 · ≥99.9% gate Architecture & Workflow Diagrams # Hand-authored, self-contained SVG views of the platform. They are theme-aware and render without any external dependency, and each links to a standalone .svg file. The 27 in-line Mermaid diagrams inside the documents are additionally rendered to SVG in the browser. Figure A — Reference Architecture (seven planes) · open SVG Figure B — The 99.9% Release-Gate Assurance Pipeline · open SVG Figure F — How Released Correctness Is Earned · open SVG Figure C — Multi-Agent Workflow (gated & evidenced) · open SVG Figure D — ASMM-Med Maturity Staircase · open SVG Figure E — Tiered Model-Fleet Routing · open SVG ← Previous Assessment Scorecard