On this page
Sit through a week of Ukrainian defense-tech meetings and you will hear the same two words in every one of them. The striking part is that the companies are usually telling the truth: the system really has been to the front, a real unit really has flown it against a real adversary. The problem arrives later, when a diligence team tries to write that down. Two products with identical claims turn out to be separated by a factor of ten in how much anyone actually knows about either — and nothing in the phrase itself tells you which one you are looking at.
The gap between the claim and the file
A performance claim in a normal technology market is checked by using the product. In defense that route is closed: you cannot run a trial of an anti-drone system against a peer adversary in your own country, which is why the claim carries so much weight and why it is examined so loosely. The diligence file ends up recording that the product is combat-proven, sourcing it to the company, and moving on to the cap table.
What the claim compresses is a wide range. At one end, a founder was told by an operator over the phone that the unit likes the system. At the other, an evaluation was scoped in advance, run over a defined period with a named mission profile, instrumented, and written up with the failures included. Both are described in a deck with the same two words. The distance between them is the entire value of the evidence.
None of this is a reason for scepticism about Ukrainian systems specifically. The underlying question of what a front establishes and what a range establishes is a real one and we have answered it separately in what battlefield testing proves. This article assumes that groundwork and asks the narrower question a buyer actually faces: given a claim in front of you, how good is the evidence, and what would make it better.
Five grades of evidence behind the same claim
The grades below are a working scale, not an industry standard. Their use is comparative: they let two products with the same sentence in the deck be placed at different points, and they tell you what the next question should be.
| Grade 1 · Anecdote | An operator or unit said the system performed well. Unwritten, unattributed, no denominator. Real information, but it cannot survive being repeated to a third party. |
|---|---|
| Grade 2 · Delivery record | Units were supplied and are in use. Establishes that a customer exists and the product is fieldable; establishes nothing about effectiveness. |
| Grade 3 · Operator feedback, recorded | Written after-action input from identifiable formations over a defined period. The first grade where a pattern rather than an impression is visible. |
| Grade 4 · Structured evaluation | Scoped in advance against a stated mission profile, measured on defined criteria, failures recorded. Methodology is disclosed and the result is reproducible in principle. |
| Grade 5 · Instrumented and independent | Grade 4, plus data captured by instrument rather than report, and run or witnessed by a party with no commercial interest in the outcome. Rare, and the only grade that transfers cleanly into another organisation's decision. |
Most claims in the market sit at grades 2 and 3. That is not a scandal — it reflects how fast systems are fielded in a war and how little of a deployment anyone has the capacity to document while it happens. It does mean that a company sitting at grade 4 or 5 is materially rarer than the uniformity of the language suggests, and that the difference is worth paying for.
The questions that move a claim up or down
Each question below is answerable without touching anything classified, and each one separates a record from an impression. The pattern of what a company can answer immediately, answer with a week of work, or not answer at all is itself the finding.
- 1How many units, over what period? A denominator turns a story into a rate. Twenty units over eight months and two units over a fortnight are different claims.
- 2Against what threat environment? Effectiveness against an unjammed target and against active electronic warfare are separate numbers. If only one exists, establish which.
- 3When? The single most informative question in this domain, and the one most often left out of the answer.
- 4What failed? An evidence pack with no failures in it has been filtered. Recorded failure modes are the strongest available signal that the rest was recorded honestly.
- 5Who recorded it, and do they sell the product? Self-reported is not disqualifying; undisclosed self-reporting is.
- 6What is the measurement, precisely? Whether a hit means a target destroyed, disabled or merely engaged decides what the percentage means, and the definitions differ between companies describing the same thing.
- 7Would the unit repeat the statement? Not always available, but where a formation will confirm sustained use, the claim moves up a grade on its own.
Validation has a shelf life, and it is short
In most technology diligence, a performance result from eighteen months ago is still broadly informative. Here it may describe a system that no longer works, without anything having changed about the system. RUSI recorded GMLRS effectiveness falling from roughly 70% in 2022 to about 30% in 2023–2024 and around 8% in 2025 as Russian electronic warfare adapted around it. The munition met its specification the whole time. What decayed was the environment, not the product.
The rate of change underneath that is the thing to internalise. CSIS documented the research, development, testing and evaluation cycle in Ukraine compressing from years to weeks; in electronic warfare, the jam and counter-jam exchange runs on a cycle of roughly three weeks. A company whose evidence is a year old is not necessarily overstating anything — it is describing a different war.
The same decay cuts the other way and is worth naming, because it is the strongest thing a well-documented company has to offer. A team that can show a system holding effectiveness across several adaptation cycles has demonstrated something no single test result can: that the engineering organisation iterates as fast as the threat. That is a durable property. A one-off result is not.
Validated in Ukraine is not qualified for your procurement
These are two different evidentiary systems and they are routinely treated as one. Operational validation answers whether the system works against a real adversary. Qualification answers whether it will behave predictably across the environmental envelope your own forces will impose, and whether it can be bought through a standard procurement process at all.
- Qualification evidence
- The laboratory and documentary layer a procurement authority requires independently of operational history: environmental testing to MIL-STD-810 or a European equivalent, electromagnetic compatibility to MIL-STD-461, a quality system at ISO 9001 or AQAP level, and — for most NATO state purchases — codification with an NCAGE code and an NSN so the item exists in the supply catalogue. None of it is produced by combat use.
For a buyer this is a scheduling question more than a quality one. A combat-validated system with no qualification package is not a worse product; it is a product that is twelve to eighteen months from being purchasable through your normal channel, and that gap belongs in the model. The underlying requirements are set out in the codification guide, and the laboratory layer in our walk-through of running an evaluation.
Where diligence-grade evidence comes from
Three channels produce evidence at grade 4 or above, and they are not interchangeable — the difference between them is who the validation was performed for and who ends up owning the record.
- State validation platforms.Ukraine's own evaluation route tests against the needs of its forces. Brave1's Battle Proven competition, held at Defense Tech Valley in Lviv on 16–17 September 2026, is a visible instance: entry requires a product deployed in active combat or projected for frontline deployment within twelve months, and Ukrainian military formations take part in the assessment. Passing that filter is a real signal; the record belongs to the platform rather than to you.
- An evaluation you commission. Scoped to your mission, measured on your criteria, documented to your standard, and owned by you. This is the only channel that produces a file you can put in front of your own board without a translation layer. Which of the two routes fits is the subject of our comparison of independent T&E and the state platform.
- The company's own operational record. Weakest as a source and still the usual starting point. It is worth more than its reputation when it is contemporaneous, includes failures and names its methodology — and worth very little when assembled retrospectively for a data room.
What belongs in the diligence file
The aim is a section a person who was not in the room can read in ten minutes and act on. Six items do that work, and the security constraint does not prevent any of them: unit identities, locations and precise dates generalise cleanly while the measurements stay intact.
- The claim as the company states it, quoted, with the date it was made.
- The grade on the scale above, with the single fact that determines it.
- Denominator and period — how many units, deployed when, for how long.
- Threat environment, described doctrinally rather than by location.
- Recorded failure modes and what was changed in response.
- The qualification gap: which of MIL-STD, quality system and codification exist today, and the time and cost to close the rest.
Read alongside the commercial and legal work rather than instead of it. Product evidence sits underneath the questions that a control acquisition or an equity investment raises about ownership, contracts and structure; a company can be immaculate on all of those and still rest its valuation on a sentence nobody has tested. Grading that sentence takes a week and is the cheapest diligence you will run.
Frequent questions
- GMLRS effectiveness falling from roughly 70% (2022) to about 30% (2023–2024) and around 8% (2025) as Russian electronic warfare adapted — RUSI, “Disrupting Russian Air Defence Production”, December 2025
- Compression of the research, development, testing and evaluation cycle from years to weeks in Ukraine — CSIS, Bondar and Bendett, 28 May 2025
- Battle Proven 2026: Brave1 startup competition at Defense Tech Valley, Lviv, 16–17 September 2026 — TechUkraine
- Battle Proven 2026 eligibility — products deployed in active combat or projected for frontline deployment within 12 months; evaluation with Ukrainian military formations — Ministry of Defence of Ukraine
- MIL-STD-810 environmental engineering considerations and laboratory tests — ASSIST / US Department of Defense
- NATO codification: NCAGE and NSN through the NSPA codification tools
Published: 1 August 2026
