Digital Engineering/AI Product Assurance/Continuous Product Assurance
Assurance that stays current while you keep shipping
Continuous Product Assurance keeps the picture current: a re-assessment cadence, a check on releases that change something consequential, and a named engineer who already knows your system.
This service is in early access, and here is what that means
Continuous Product Assurance is running with a deliberately small number of teams. It is not a finished, packaged, price-listed service, and we will not pretend otherwise.
The reason is arithmetic. A cadence is only useful once calibrated: how fast a codebase of a given size drifts, which release types are worth gating and which are noise, how long a monthly pass takes. We are building that baseline against real systems rather than inventing it.
What the early group gets
Influence over how the cadence and the gate are designed, generous engineer time while we calibrate, and terms reflecting that you are helping build it.
What we ask
Tolerance for a service still being shaped, honest feedback when a check proves worthless, and permission to use anonymised drift data, never your name, code or findings.
How to join
Talk to us. The one requirement is a current assessment against the framework: ours, or one scheduled as the starting point.
A findings register starts decaying the moment it is delivered
The register we hand over is accurate about a specific commit on a specific date, and stops being fully accurate the next time anyone merges anything. That is not a weakness of the assessment; it is what happens when a document describes a system still being built.
The decay arrives through a few identifiable mechanisms, every one a normal part of shipping software.
New dependencies.
Every package brings its own transitive tree, licence position and future disclosure. One added on a Tuesday can carry twenty things nobody chose and nobody inventoried.
New routes.
Each endpoint is an authorisation decision: made deliberately, made by pattern-matching against the route beside it, or not made at all. The assessment could only judge routes that existed when it ran.
New data flows.
A new integration, export, webhook or analytics call is a fresh path along which customer data moves, and the assessment's classification work did not cover it.
Staff change.
Findings are often about knowledge rather than code. When the person who understood how invoices are calculated leaves, a Ready dimension becomes Needs Attention without a line changing.
Tooling change.
A model upgrade, a new IDE or a revision to your agent rules changes what gets generated. Code written afterwards need not resemble the code the assessment read.
And remediation changes the system.
Closing Critical findings moves boundaries and touches shared code, so remaining findings were classified against a system the fix has since altered.
Run those for two quarters and the register's useful half-life has passed. Teams discover this at the worst moment, when a customer, an insurer or an investor asks for the current position and the only document carries last quarter's date.
What continuous assurance actually is
Four things running together. None is a new framework: the measurement stays comparable to the assessment you have, and re-applies the published methodology.
A re-assessment cadence against the same six dimensions.
The Noseberry Product Assurance Framework, covering Code Health, Application Security, Product Architecture, Scalability & Performance, Cloud & Delivery and Engineering Governance, is re-applied on a defined rhythm, using the published evidence rules. Each carries a dated status of Critical, Needs Attention or Ready, so consecutive assessments compare rather than merely differ in wording.
A gate on releases that change something consequential.
Not every release. A defined trigger set: authentication and authorisation changes, data model migrations, new external dependencies, anything altering how payment or personal data is handled, infrastructure or secrets configuration. Everything else ships without us.
Drift detection between full assessments.
Automated checks run continuously against the repository and pipeline, watching for the mechanisms above: packages entering the tree, routes appearing without a matching authorisation pattern, new outbound data destinations, coverage falling on modules whose findings were closed, configuration diverging from the last recorded state.
A named engineer who already knows the system.
The same person, month after month, who read your codebase during the assessment and does not need re-briefing. An engineer with six months of context spots a bad change while a stranger is still working out what the module does.
What the automation does, and what it does not
What the automation catches
Automation makes daily coverage affordable. It watches every commit and catches mechanical drift: a package appeared, a secret was committed, a config value diverged, a route lacks a guard its neighbours have.
What only a human catches
It cannot tell you that the new export endpoint returns a field three enterprise contracts say you will not expose, that the migration was correct and the backfill was not, or that two teams have built competing permission logic. That is why a human review sits on the cadence, and why we do not describe the automation as the service.
Cadence options
Three rhythms, chosen against how fast you ship and what your customers require. Most run the monthly pass plus release gating, adding the quarterly full re-assessment once a customer or investor asks for a dated position.
Per-release sign-off.
A review of a specific release before it goes out, triggered only by the change types above. It covers what that release changes and how it interacts with what exists. It does not cover the rest of the system, and a clear verdict says nothing about the parts of your product that release did not touch.
Monthly dimension re-assessment.
A pass across all six dimensions using automated evidence plus a half-day of engineer review, producing dated statuses and a revised findings register. It covers direction of travel and four weeks of drift. It lacks the depth of a full assessment: no fresh architecture walkthrough, no manual review of every module.
Quarterly full re-assessment.
The complete assessment run again: automated pass, engineering review across the whole system, reclassification of every open finding, refreshed statuses with evidence. It produces a defensible current position and a genuine trend line. It does not cover remediation; findings still have to be fixed.
These stack rather than substitute. Gating without a periodic full pass tells you each change was reasonable while the accumulation moves a dimension backwards. A quarterly pass without gating tells you, twelve weeks late, that something went in that should not have.
Release gating in practice, including how you overrule us
The objection is always the same and it is fair: you built with AI tooling partly to move quickly, and an external reviewer standing between your team and production sounds like the wrong trade. Here is the mechanism.
What we return: one of three verdicts
The release is consistent with the framework. It ships with nothing further from us.
The finding is attached, the release can go out, and the risk is recorded with an owner and a review date.
A finding should be resolved first. Where one exists, we give the smallest change that converts a Hold into a Clear.
Turnaround is measured in working hours and written into the subscription. The review runs against the pull request, not the deployment, so it happens while the work is in flight rather than when someone wants to press the button.
What triggers the gate
A written trigger set agreed at onboarding and encoded in your pipeline, not a judgement call made in the moment. Authorisation and authentication changes, data model migrations, new external dependencies, payment or personal data handling, infrastructure and secrets configuration. In a typical month that catches between one release in five and one in twenty. Everything outside the set ships with no involvement from us.
Who decides
You do. We hold no veto and would not want one. We are not accountable for your commercial calendar and cannot weigh a Hold against a customer commitment. Your named decision-maker, agreed at onboarding, decides what ships.
How an override works
One line in the release record: who overrode, and why. The finding moves into the register as an accepted risk with an owner and a review date, so it stays visible. We do not argue about overrides, and an overridden Hold is not held against the dimension status once the underlying risk is closed.
And if we are slow, you ship
If a review is not returned inside the agreed window, the release proceeds by default. The default state of this service is proceed, not stop. An assurance process that can quietly become a release blocker gets routed around within a month, and then it protects nothing.
Who this is for
Teams that have completed an assessment and want to hold the position.
You spent real money finding out where you stood and real effort closing the Critical findings. Watching that decay over two quarters is the commonest waste in this category.
Teams selling to enterprise customers with ongoing obligations.
Your contracts do not say “was secure in March”. Procurement routinely asks for a current position or evidence of continuous practice, and a document dated nine months ago is a weak answer.
Teams shipping quickly with heavy AI assistance.
You want a check that scales with the volume of generated code rather than a review process that becomes the bottleneck you were avoiding.
Teams heading towards a fundraise or a sale.
Diligence rewards a trend more than a snapshot. Three dated assessments showing improvement are a different conversation from one run three weeks before the data room opened.
Teams where one engineer holds all the context.
A named external engineer with standing knowledge of the system is a real continuity control, and cheaper than discovering what that person knew once they have gone.
How it works month to month
Onboarding
week one
We establish the baseline. If you have a recent assessment we use it; if not, a full assessment runs first, because there is otherwise nothing to detect drift against. We agree the trigger set, the turnaround window, the decision-maker on your side and the engineer on ours. Read-only repository access and read visibility of the pipeline go to named reviewers. We take no write access and no production credentials.
The automated layer
continuously
Dependency, secrets, static analysis and configuration checks run against every merge, tuned to your stack rather than left at defaults. Anything meeting the Critical threshold reaches your team the same day; everything else accumulates into the monthly pass rather than generating alerts nobody reads.
The gate
on release
Triggered only by the agreed change types, reviewed against the pull request, returned inside the window, verdict recorded, your team decides.
The review pass
monthly
Your engineer reviews what changed rather than re-reading the whole system: the diff, the new surfaces, the drift signals, the findings meant to close. Statuses are updated with dates, the register is revised, and a 45-minute call covers what moved.
The full re-assessment
quarterly
Every open finding is reclassified against the current system and the assurance statement is reissued with a new date and a new set of statuses.
The baseline review
annually
We check whether the cadence is still right. Teams that have stabilised often move to a lighter rhythm.
What you receive
- A current assurance statement, reissued each quarter, describing what was assessed, when, against which dimensions and what status each holds, written so a customer's security team or an investor's adviser can read it without a call.
- A live findings register, dated at every revision, showing open findings by class, what closed, what was newly identified, and what was accepted as a risk with an owner and a review date.
- A monthly written update covering what changed, which dimensions moved and why, which drift signals fired, and the two or three things worth doing next.
- A remediation trend, showing open findings by class over time and the median time to close each class, often more convincing to an enterprise buyer than the current status.
- Release verdicts on record, each dated, with the change reviewed, the verdict given and any override recorded with its reason: an evidence trail of engineering practice rather than an assertion of it.
- A quarterly assessment report covering all six dimensions with the evidence behind each status, in the original engagement's format so documents sit side by side.
- A named engineer, reachable during agreed hours, who has read your system and needs no re-explanation.
The second assessment is worth more than the first
The first assessment tells you where you are, which is valuable exactly once. The second and third are where the useful information lives, because what a sophisticated reader wants is not a status but a direction.
Consider two documents put in front of an enterprise security team. The first says Application Security is Ready. The second says Application Security was Needs Attention in Q1 with nine open findings across authorisation and secrets handling, is Ready in Q3 with those closed and evidenced, and that the median time to close a High finding was eleven days. The first is a claim. The second is an argument, and it survives questioning.
Investors read it the same way. A single clean assessment run shortly before a data room opens invites the obvious question of what came before. A trend spanning three quarters answers a different one: does this team find problems and close them, and how quickly. That is a statement about engineering capability, what technical due diligence exists to establish.
How trend works before numeric scoring exists. Numeric composite scoring is held until the methodology page is published in full, so the trend is built from what is already defensible: the dated status of each dimension per period, open findings by class over time, findings closed per period, and median time to close by class. Plot those and you have a real trend line with no invented precision. When scoring arrives it will be applied to historic assessments too, so early subscribers get a back-dated series rather than starting at zero.
How this fits with the rest of the cluster
This is the layer that comes after the one-off work across AI Product Assurance, and replaces none of it.
AI Code Audit is the entry point and, in practice, the prerequisite: it produces the baseline this service maintains. Without one there is nothing for drift to be measured against.
Vibe Code Cleanup and Production Readiness & Hardening are the remediation engagements. Continuous assurance identifies and tracks; those two execute. Work too large for the monthly rhythm is scoped separately rather than quietly consuming the retainer.
AI Application Security Audit goes deeper than any recurring pass can. Teams under real security pressure run one annually alongside the subscription, and its findings feed the same register.
Software Technical Due Diligence draws on whatever history exists. A subscriber arriving with three dated assessments needs far less preparation than a team starting from nothing.
AI Development Governance defines how your team is meant to work with AI tooling. This shows whether those rules are followed in the code, which governance documents rarely evidence.
What this does not do
It is not incident response.
If your product is down at 2am, this subscription does not answer. There is no on-call rota attached and no availability commitment of any kind. Incident support is a separate arrangement.
It is not feature work.
Engineer time goes on assessment, review and gate decisions, not a pool of development hours. The moment assurance time becomes delivery capacity, the assurance stops happening.
It is not a certification and it is not a guarantee.
The assurance statement records what was assessed, when, and against what. We identify gaps relevant to OWASP ASVS, the OWASP Top 10 and NIST SSDF, and the work supports readiness for a SOC 2 or ISO 27001 process. We do not certify compliance against any of them, and no engineering review can.
It is not a penetration test.
Adversarial testing is a different discipline. We will tell you when one is warranted and read the report, but we do not run it here.
It does not catch everything between reviews.
On a monthly cadence, up to four weeks can pass before a human sees a change that did not trigger the gate. That is a real limitation and the reason the gate exists.
It does not replace your own testing.
Your CI, test suite and code review remain yours; this sits above them and substitutes for neither.
And it does not stop you shipping.
Overrule a Hold and the release goes out, the decision is recorded, and the consequence is yours.
Why Noseberry for this
Continuity through the same engineer.
The value compounds through the person, not the process. Your engineer read the codebase during the assessment and carries that context forward. Rotating reviewers would be cheaper to staff and a materially worse service.
One framework across every service and every period.
Your Q1 assessment, your Q3 re-assessment, your security audit and your diligence pack all report against the same six dimensions. Nothing has to be reconciled, and the trend is a trend rather than a comparison of different measurements.
We publish the method.
The evidence rules behind every status are on the methodology page. Recurring assessment only means something if the measurement is stable, and stability is only checkable if the rules are public.
We tell you when the cadence is wrong.
Some subscribers should be on a lighter rhythm than they pay for, and we raise it at the annual review.
We are candid about maturity.
This page says the service is in early access because it is. A supplier who overstates before launch will overstate a finding after it.
Frequently asked questions
An audit is a diagnosis with a date on it. This is the same measurement repeated on a rhythm, plus a check on releases that change something consequential, plus continuous drift detection, plus a named engineer who does not need re-briefing. The framework and evidence rules are identical, which is the point: comparability is what produces a trend. If you need your position once, buy the audit. If customers or investors will ask again in six months, repeating the audit from scratch is the expensive route.
It is designed not to. The gate fires only on an agreed trigger set: authorisation changes, migrations, new dependencies, payment or personal data handling, infrastructure and secrets configuration. That is a small minority of releases in most months. Reviews run against the pull request rather than the deployment, turnaround is a contracted window measured in working hours, and if we miss it the release proceeds by default. We hold no veto: your named decision-maker decides, and an override takes one line to record.
Quote-led, deliberately, while the service is in early access. Price is driven by codebase size, release frequency, how many of your releases hit the trigger set, and which cadence you run. The shape of the conversation is a monthly subscription with a defined engineer-time allocation, a stated gate turnaround window, and the quarterly re-assessment either included or priced separately by tier. There is no setup fee beyond the baseline assessment, and early-group terms reflect that you are helping us calibrate.
Not really, and we would rather say so. Drift detection needs something to measure against. Without a baseline, the first three months are spent discovering what was already there, which is an audit run slowly at higher cost. If you have not had one, the subscription starts with a full assessment as month zero. If you hold a recent assessment from another supplier, send it: where it covers comparable ground we will map it onto the six dimensions rather than charge you to repeat it.
Yes, with a limited number of teams, and we would rather be plain than imply a maturity the service does not have. What exists today: the framework, the assessment method, the automated drift checks, the gate mechanism and named engineers. What is still being calibrated: cadence timings, the trigger sets that prove worthwhile per stack profile, and standardised pricing tiers. We are taking a few more teams while that baseline is built. If you would rather wait until it is packaged, say so and we will come back to you.
You are told who it is before you commit and you meet them at onboarding. Continuity is the product, so handover is treated as a real risk rather than an administrative detail. A second engineer is briefed on your system within the first quarter and stays lightly current through the monthly passes, so a handover is a transfer rather than a restart. If your named engineer changes you are told in advance, and the incoming engineer runs the next full pass alongside the outgoing one.
Yes, that is largely what it is for. It is written to be read outside your organisation and states plainly what was assessed, when, against which dimensions, and what status each holds. Be careful how it is described: it supports readiness for a security review or a diligence process, it does not certify anything, and it is not a guarantee about your product. We will produce a version aimed at a non-technical reader on request. We share nothing with any third party without your written instruction.
You hear about it the day we find it. Critical findings from the automated layer are escalated to your named engineer and your team immediately rather than held for the monthly pass. A finding that waits three weeks for a scheduled document has failed at its job. Where the fix is small and within scope we describe exactly what to change. Where it is a body of work, we scope it honestly rather than absorbing it silently into the retainer and delivering half of it.
Read-only in both cases. Repository read access and read visibility of pipeline and configuration state cover everything described here, including the gate: a release verdict is a comment on a pull request, not a merge we perform. We do not need production database access or customer data. If a remediation engagement follows, write access is granted separately, to named engineers, against agreed branches, and it ends when that work does. Where your policy demands it, scanning runs inside your own infrastructure and only results leave.
Yes, on a notice period agreed at the start, and you leave with everything: the full findings register, every assurance statement and every quarterly report, in a portable format. There is no lock-in through data retention and nothing in the arrangement that only works while you keep paying. Teams that pause and return are common; a heavy build quarter is a reasonable time to step back, though a longer gap means the first pass back is closer to a full re-assessment.
Join the early group
We are working with a small number of teams while the baseline is established, and there is room for a few more. Send us your most recent assessment, or the repository if you have not had one, and a note on how often you ship. We will come back within two working days with a proposed cadence, a trigger set, an engineer, and an honest view of whether you need this yet.
Related resources
Go deeper with our thinking
From Spreadsheets to a Scalable Platform – building Canada’s Most Tenant- First Coliving System
Read case studyHow to Build a SaaS Product
Read guideThe MVP Development Guide
Read guideBuilding a Digital Engineering Roadmap That Aligns With Business Goals
Read articleEngineering IoT Products: Key Challenges and Best Practices
Read articleNot sure where you stand?
Take a free two-minute readiness scorecard built for your industry.
