tech demo · public sector
Anticipating the discount and the empty tenders in public procurement
We sealed the rules before looking at 2024, then ran the exam over 120,494 real tenders. On the award discount we beat the simple rule; on the empty tenders the simple rule catches more cases and the model ranks the risk better. With its misses and its caveats.
Spain publishes hundreds of thousands of public tenders every year: contracts through which a public body buys something, from cleaning a building to building a road, with companies bidding to do the work. Almost all of it is public, and it leaves a huge trail of data.
Two things in that trail are worth money before they happen. One is the discount: how much cheaper than the starting budget the contract is finally awarded. The other is the tender coming out empty, meaning no one shows up or no bid is good enough and the contract closes without an award. One in ten comes out empty. Knowing the likely discount, or the risk of a contract going unanswered, is useful both to the company deciding whether to bid and to the authority that wants its tender not to fail.
The conclusion, up front: the protocol of hiding the future works and is verifiable. The model estimates the discount better than a simple rule, and on the risk of a tender coming out empty it wins on the quality of the probability though not on coverage. It also knows when to stay quiet: it abstains on 11% of cases.
status: exam run on 15 July 2026
The order matters and we kept it: first we sealed the rules, then we ran the model. The formulas, the thresholds and the cut-off date were written down and sealed on 15 July 2026, before a single prediction was computed. The seal is cryptographic (OpenTimestamps, anchored in Bitcoin), so nobody, not even us, could touch them afterwards.
With the rules closed we ran the exam: we trained only on what was resolved up to 31 December 2023 and predicted over 120,494 lots published in 2024 and resolved in 2024. The numbers come from there and are published exactly as they came out, with the wins and the misses.
| Year | Published | Not yet resolved | Resolved | Awarded | No-bids | Withdrawals and waivers |
|---|---|---|---|---|---|---|
| 2022 | 231,733 | 27,938 | 203,795 | 178,158 | 22,614 | 3,023 |
| 2023 | 251,028 | 27,679 | 223,349 | 198,411 | 21,835 | 3,103 |
| 2024 | 261,674 | 28,375 | 233,299 | 209,373 | 20,476 | 3,450 |
| Year | The body's regular | Another with track record | A newcomer | No identifiable tax ID |
|---|---|---|---|---|
| 2022 | 8,377 | 117,773 | 42,135 | 9,873 |
| 2023 | 9,882 | 146,381 | 31,896 | 10,252 |
| 2024 | 10,768 | 165,502 | 23,347 | 9,756 |
What we already have verified
Before promising any prediction, the boring work: download the official history, parse it, remove duplicates and cancelled records, and check by hand that the numbers match the public record of each contract. That is done.
744,435
award rows, at lot granularity, in the final table
560,462
live records, after dropping 45,017 cancelled ones
8.8 to 11.1%
come out empty, depending on the year (the share is falling)
6%
median award discount (mean 16%, long tail)
30%
awarded with no discount: at the starting price, many with a single bidder
34%
of the rows belong to multi-lot contracts: treating them as one would be an error
The data was cross-checked by hand against the public record of four real contracts, and all four match 100% on budget, award amount, awardee, number of bidders and outcome. Counts, breakdown by year and by sector, and a spot check of the values: all three layers pass.
How we did it
The method is the same as in all our demos: a backtest with the future hidden. We pick a cut-off date, train the model only on tenders before that date, seal the rules, and only then uncover what happened afterwards to see whether the prediction was right. Sealing means writing the formulas and the thresholds down before looking, so they cannot be tweaked to make the result look good.
The biggest technical risk here has a name: future leakage. The official history is a file that keeps updating, and the same record appears several times as it changes state. If the cut-off is not made by the date of each notice, information from the future slips in and the backtest lies. That is why we keep the date of each publication and of each award separately, and the cut-off was made by those, never by the year of the file. That there was no cheating with the future we checked three ways: by design (every historical statistic only counts what was resolved before the publication day), with automated tests, and by recomputing a real case by hand until the number matched the machine exactly.
All three deliveries are run. The results, below.
first delivery · done
The discount and the empty tenders
We anticipated how much an award would drop against its budget, and the risk of the contract coming out empty. It was the most direct: both figures are clean and computable. The result, right below.
second delivery · done
Which tenders fit a company
A radar that, from what each company already won, flags the new tenders that fit it. It beats the simple rule by twice the required margin. The result, further down.
third delivery · done
Who wins it, in aggregate
Anticipating which company wins each contract. The uncomfortable, published result: inertia beats the model, and the real ceiling is far lower than we expected. Further down.
The verifiable result
The exam ran over 120,494 lots from 2024. The verdict is twofold and told in full. Everything is compared against a simple rule, the one anyone could apply without a model: for the discount, always predict that sector's typical drop; for the empty tenders, the rate of empty tenders that sector and that kind of contract carried from before. Beating that rule is the condition for success, not a pretty number.
5.7 / 6.4 pp
overall typical error, in percentage points (model / simple rule): the model gets closer
1.2 / 4.7 pp
on the third that is awarded with no discount, the model almost nails the zero (model / simple rule)
8 / 8 pp
in real competition, a technical tie: getting the exact discount there is hard (model / simple rule)
How to read it: the typical error is how many percentage points, in a normal case, we are off from the discount that actually happened (so lower is better). The model wins overall and wins comfortably where there is learnable signal: it recognises the third of contracts that close at the starting price and predicts near zero, while the simple rule always drags several points of error there. In real competition, predicting the concrete discount is genuinely hard and we barely improve on the simple rule. The big win comes from the easy case. And against the yardstick we sealed before looking: 69.9% of our predictions fall within the 10-point band fixed in the pre-registration (the simple rule, 68.7%).
9.9%
of the empty tenders the model catches as a yes/no light, at equal precision
15.5%
the simple rule catches: more coverage than the model, at equal precision
12%
when the model says 12% risk, 12% come out empty: the probability is honest
How to read it: as a yes/no light at equal precision, the simple rule catches more empty tenders than the model. Where the model does win is on the quality of the probability: it ranks the risk better and its percentages say what they are (when it warns of a 12% risk, 12% come out empty). It is useful for prioritising where to look first, not for a dumb light.
11%
of the lots the system says "I don't know" and does not score, mostly in bodies with little prior history. It falls within the 10-20% band we fixed before looking. On what it does not abstain from, the numbers above hold. A system that knows when it does not know gives more confidence than one that always answers.
A look at the misses
Two of the biggest misses, with their record:
- False alarms. Three construction jobs from the public company Tragsa, with similar budgets, got a 33.6% risk of coming out empty for falling into a sector-and-procedure crossing with a high past rate. All three were awarded and signed without trouble: the sector's risk does not override a recurring buyer who almost always closes its jobs.
- Empty tenders nobody anticipated. Three records the model gave a 0.5% risk ended up empty: a port authority service and two supplies, one to defence and one to a technology institute in Valencia. They fell into sectors with a very low past rate of empty tenders; the model trusted the history and the history failed. It is the kind of surprise no prior data saw coming.
The radar: which tenders fit a company
The second delivery builds a radar: for each company with enough track record, it ranks each month's new tenders by how well they fit what that company already used to win, and shows it the top ten. The measure is recall@10: of the tenders the company actually won that month, how many were in that list of ten. With the future hidden as before, the radar beats a simple rule ("more of the same sector and the same buyer") by twice the margin we fixed in advance.
+9.2 pp
the radar's margin over the simple rule on recall@10 (30.4% against 21.2%); the sealed bar asked for 5 pp
47%
the radar's hit rate on companies with little history (5 to 9 contracts won), against the simple rule's 31%: with a thin profile, reading the text rescues more
20.9%
of the companies with any history are the ones the radar trusts and speaks about (11,123 of 53,314); on the rest it abstains
How to read it: the radar gains most where the company's profile is thinnest, which is exactly where a rule of "more of the same" has least to work with. The big caveat goes up front: we only know what each company won, not everything it bid on, so a tender that fit it but lost counts here as a miss. The recall we measure is a floor, the real radar hits more than this exam credits it with. The embeddings that read the contract's object are an open-weights model (Qwen3) run on a local machine, with no API calls.
Who works with whom
The same history answers another question: crossing the twelve bodies that award the most with their awarded suppliers, do patterns show up? They do, and they are sectoral and clean. It is not a ranking or an accusation: it is the picture of the split, exactly as it stands in the public records.
| Body | Awards | Main named awarded suppliers |
|---|---|---|
| TRAGSA | 16,067 | dispersed split: almost all (16,053) outside the top 15 |
| Valladolid Univ. Hospital | 11,916 | IBERDROLA 143, Becton Dickinson 143, Medtronic 139, B. Braun 115 |
| State Central Purchasing | 11,632 | CANON 620, IBERDROLA 450, INETUM 425, ENDESA 376, CEPSA 335 |
| Renfe Eng. & Maint. | 8,216 | Castro Perero 514, I2M 432, Faiveley 359 |
| Paradores | 6,811 | Frutas y Verduras Massanassa 319 |
| Insular Univ. Hospital (Canary Is.) | 5,780 | B. Braun 108, Medtronic 76, Palex 53 |
| Cantabria Health Service | 4,442 | Becton Dickinson 181, B. Braun 150, Medtronic 148, Palex 117 |
| ADIF (Chair) | 3,582 | dispersed split: almost all outside the top 15 |
| Asturias Care Homes | 3,222 | INETUM 19, Medtronic 18, Telefónica Sol. 17 |
| FREMAP | 3,059 | B. Braun 10, INETUM 8, Medtronic 8 |
| Murcia Health Service | 2,809 | Medtronic 100, Palex 58, B. Braun 14 |
| RTVE (procurement) | 2,765 | Telefónica Sol. 18, INETUM 11, IBERDROLA 9 |
Who wins it: inertia beats the model
The third delivery is the most ambitious and the one that comes out worst for an easy headline, which is why we publish it all the same. The question was whether you can anticipate which company will win a tender. The dumbest rule there is, "it goes to whoever has won it most before with that same buyer in that same sector", gets the exact name right 12.1% of the time. Our model, with the future hidden, falls short: 9.5%. It does not beat inertia, and the pre-registration already said that, if it did not, the finding would be inertia itself.
12.1 / 9.5%
getting the winner's exact name right first time (inertia / the model): here the simple rule wins, and we publish it
59.2%
is the reachable ceiling: four in ten tenders are won by a company new to that body, impossible with this method. Our pre-registration estimated it at 95.6%, inflated by measuring without hiding the future
16.3%
of cases the system abstains, when there are too many candidates or too little confidence; within the 10-20% band fixed before looking
How to read it: getting it right first time is top-1; if we give three names (top-3), inertia puts the winner in 22.8% of the time and the model 20.1%; in a list of ten they tie (35.8% against 35.6%). The model beats the weak baselines, so it has learned something real, but mixing signals it sometimes overthinks and puts a bigger company ahead of the one that actually repeats. The useful reading for a company is not guessing the name, it is knowing where there is no usual supplier: that is where a newcomer really has a chance.
The caveats, from the start
- It is not 100% of public spending. We work with the State's contracting profiles, without the minor contracts or the own platforms of some regions. It is a large and representative slice, but not the whole, and we will state the exact scope.
- This is not virgin ground, and we cite it. Researchers at the University of Oviedo already published very similar estimators on Spanish procurement data: an award price estimator and a bidder recommender. The ML of this is already invented. Ours is not the model, it is the protocol of sealing the rules before looking and showing the prediction against what actually happened.
- The tail of the discount is dirty in the raw. There are records with tiny budgets that produce absurd percentages; for the statistics you have to bound it to the plausible range, and when the discount cannot be computed it is left empty, never a false zero. Even so, the odd artefact slips through: the three worst discount errors are awards with a "100% discount" (an amount near zero against the budget), almost certainly not real discounts but noise in the data, amounts loaded wrong. The model neither can nor should predict that, and we say so.
- 18% of the awards fall outside the cut-off. They carry no award date in the raw data, so they cannot be placed in time without inventing one, and we do not invent it. The exam still holds 102,867 awards, plenty to measure on.
- The number of bidders is not always present. It is in just under nine out of ten resolved rows. In any case it is not used to predict: it is information that comes after publication.
What about your company?
The same reading of the history works to budget a job or a service by reading what similar ones cost instead of guessing, to know in advance which calls, offers or internal tenders will go unanswered, or to estimate how much an auction will close at before entering it.
If one of these experiments reminds you of a problem of yours, a history nobody uses, decisions that depend on a veteran's hunch, public data nobody cross-references, write to us and we'll talk it through. No fluff: we'll tell you whether it can be done.
Built with Python (streaming lxml over the official history), Parquet storage and classical machine learning (scikit-learn), with a fixed seed and pinned versions so the experiment reproduces from scratch. No language model makes the decision: it is statistics and classical models over public data.
Data and thanks
This experiment exists because Spanish public procurement is published openly. Thank you.
- Public Sector Procurement Platform (Ministry of Finance): the open data of the contracting profiles.
No affiliation with any authority or company. The names of awardees that may appear are the public ones from the procurement records themselves.
Can you use this? Yes. Our computations, the funnel and the text on this page are published under a CC BY 4.0 licence: use, share or analyse them freely, citing the source, Team Banzai (team-banzai.com). The public procurement data keeps the terms of its official source.