Skip to content
team banzai

tech demo · public sector

Anticipating the discount and the empty tenders in public procurement

We sealed the rules before looking at 2024, then ran the exam over 120,494 real tenders. On the award discount we beat the simple rule; on the empty tenders the simple rule catches more cases and the model ranks the risk better. With its misses and its caveats.

Spain publishes hundreds of thousands of public tenders every year: contracts through which a public body buys something, from cleaning a building to building a road, with companies bidding to do the work. Almost all of it is public, and it leaves a huge trail of data.

Two things in that trail are worth money before they happen. One is the discount: how much cheaper than the starting budget the contract is finally awarded. The other is the tender coming out empty, meaning no one shows up or no bid is good enough and the contract closes without an award. One in ten comes out empty. Knowing the likely discount, or the risk of a contract going unanswered, is useful both to the company deciding whether to bid and to the authority that wants its tender not to fail.

The conclusion, up front: the protocol of hiding the future works and is verifiable. The model estimates the discount better than a simple rule, and on the risk of a tender coming out empty it wins on the quality of the probability though not on coverage. It also knows when to stay quiet: it abstains on 11% of cases.

status: exam run on 15 July 2026

The order matters and we kept it: first we sealed the rules, then we ran the model. The formulas, the thresholds and the cut-off date were written down and sealed on 15 July 2026, before a single prediction was computed. The seal is cryptographic (OpenTimestamps, anchored in Bitcoin), so nobody, not even us, could touch them afterwards.

With the rules closed we ran the exam: we trained only on what was resolved up to 31 December 2023 and predicted over 120,494 lots published in 2024 and resolved in 2024. The numbers come from there and are published exactly as they came out, with the wins and the misses.

the public procurement funnel state profiles · per lot · verified data
feed year

Flow figures, by year (rows per lot)
Year Published Not yet resolved Resolved Awarded No-bids Withdrawals and waivers
2022231,73327,938203,795178,15822,6143,023
2023251,02827,679223,349198,41121,8353,103
2024261,67428,375233,299209,37320,4763,450
How the awarded split: who wins them
Year The body's regular Another with track record A newcomer No identifiable tax ID
20228,377117,77342,1359,873
20239,882146,38131,89610,252
202410,768165,50223,3479,756
How to read it: each contract can be split into several lots, and we count per lot, so the figure on the left is the number of published rows, not of contracts. From there the flow branches in two: most of it reaches an outcome and another branch is still unresolved. Of what does have an outcome, almost all is awarded, a portion in amber comes out empty and a minority are withdrawals and waivers. Each arrow carries its figure and, on hover, its percentage. The trend of empty tenders is downward: from 11.1% in 2022 to 8.8% in 2024. And one more layer: the awarded open up by who wins them. About 1 in 20 goes to the body's regular supplier, the one that had won the most contracts from it before, computed looking only at the past of each case. Most go to another firm that had already won something somewhere; the rest are newcomers with no prior trace, or awards with no winner tax ID that are not forced into any category. The regular's weight rises year on year (4.7% in 2022, 5.1% in 2024), but that is not only the market closing up: in 2022 there is barely any accumulated past, so many bodies still have no identifiable regular and their winner falls into the other buckets. The more history we know, the sharper the usual one shows. The reliable figure for the steady state is 2024's, and it is trending up.

What we already have verified

Before promising any prediction, the boring work: download the official history, parse it, remove duplicates and cancelled records, and check by hand that the numbers match the public record of each contract. That is done.

the data record feed 2022-2024 · contracting profiles, minor contracts excluded

744,435

award rows, at lot granularity, in the final table

560,462

live records, after dropping 45,017 cancelled ones

8.8 to 11.1%

come out empty, depending on the year (the share is falling)

6%

median award discount (mean 16%, long tail)

30%

awarded with no discount: at the starting price, many with a single bidder

34%

of the rows belong to multi-lot contracts: treating them as one would be an error

The data was cross-checked by hand against the public record of four real contracts, and all four match 100% on budget, award amount, awardee, number of bidders and outcome. Counts, breakdown by year and by sector, and a spot check of the values: all three layers pass.

How we did it

The method is the same as in all our demos: a backtest with the future hidden. We pick a cut-off date, train the model only on tenders before that date, seal the rules, and only then uncover what happened afterwards to see whether the prediction was right. Sealing means writing the formulas and the thresholds down before looking, so they cannot be tweaked to make the result look good.

The biggest technical risk here has a name: future leakage. The official history is a file that keeps updating, and the same record appears several times as it changes state. If the cut-off is not made by the date of each notice, information from the future slips in and the backtest lies. That is why we keep the date of each publication and of each award separately, and the cut-off was made by those, never by the year of the file. That there was no cheating with the future we checked three ways: by design (every historical statistic only counts what was resolved before the publication day), with automated tests, and by recomputing a real case by hand until the number matched the machine exactly.

All three deliveries are run. The results, below.

first delivery · done

The discount and the empty tenders

We anticipated how much an award would drop against its budget, and the risk of the contract coming out empty. It was the most direct: both figures are clean and computable. The result, right below.

second delivery · done

Which tenders fit a company

A radar that, from what each company already won, flags the new tenders that fit it. It beats the simple rule by twice the required margin. The result, further down.

third delivery · done

Who wins it, in aggregate

Anticipating which company wins each contract. The uncomfortable, published result: inertia beats the model, and the real ceiling is far lower than we expected. Further down.

The verifiable result

The exam ran over 120,494 lots from 2024. The verdict is twofold and told in full. Everything is compared against a simple rule, the one anyone could apply without a model: for the discount, always predict that sector's typical drop; for the empty tenders, the rate of empty tenders that sector and that kind of contract carried from before. Beating that rule is the condition for success, not a pretty number.

the award discount: we win typical error in points · lower is better · model vs simple rule

5.7 / 6.4 pp

overall typical error, in percentage points (model / simple rule): the model gets closer

1.2 / 4.7 pp

on the third that is awarded with no discount, the model almost nails the zero (model / simple rule)

8 / 8 pp

in real competition, a technical tie: getting the exact discount there is hard (model / simple rule)

How to read it: the typical error is how many percentage points, in a normal case, we are off from the discount that actually happened (so lower is better). The model wins overall and wins comfortably where there is learnable signal: it recognises the third of contracts that close at the starting price and predicts near zero, while the simple rule always drags several points of error there. In real competition, predicting the concrete discount is genuinely hard and we barely improve on the simple rule. The big win comes from the easy case. And against the yardstick we sealed before looking: 69.9% of our predictions fall within the 10-point band fixed in the pre-registration (the simple rule, 68.7%).

the empty tenders: depends what for as a yes/no light and as a risk meter · model vs simple rule

9.9%

of the empty tenders the model catches as a yes/no light, at equal precision

15.5%

the simple rule catches: more coverage than the model, at equal precision

12%

when the model says 12% risk, 12% come out empty: the probability is honest

How to read it: as a yes/no light at equal precision, the simple rule catches more empty tenders than the model. Where the model does win is on the quality of the probability: it ranks the risk better and its percentages say what they are (when it warns of a 12% risk, 12% come out empty). It is useful for prioritising where to look first, not for a dumb light.

knowing when to stay quiet abstention · sealed band 10-20%

11%

of the lots the system says "I don't know" and does not score, mostly in bodies with little prior history. It falls within the 10-20% band we fixed before looking. On what it does not abstain from, the numbers above hold. A system that knows when it does not know gives more confidence than one that always answers.

A look at the misses

Two of the biggest misses, with their record:

The radar: which tenders fit a company

The second delivery builds a radar: for each company with enough track record, it ranks each month's new tenders by how well they fit what that company already used to win, and shows it the top ten. The measure is recall@10: of the tenders the company actually won that month, how many were in that list of ten. With the future hidden as before, the radar beats a simple rule ("more of the same sector and the same buyer") by twice the margin we fixed in advance.

the radar: wins by twice the margin recall@10 · what share of the wins was in the list of ten · model vs simple rule

+9.2 pp

the radar's margin over the simple rule on recall@10 (30.4% against 21.2%); the sealed bar asked for 5 pp

47%

the radar's hit rate on companies with little history (5 to 9 contracts won), against the simple rule's 31%: with a thin profile, reading the text rescues more

20.9%

of the companies with any history are the ones the radar trusts and speaks about (11,123 of 53,314); on the rest it abstains

How to read it: the radar gains most where the company's profile is thinnest, which is exactly where a rule of "more of the same" has least to work with. The big caveat goes up front: we only know what each company won, not everything it bid on, so a tender that fit it but lost counts here as a miss. The recall we measure is a floor, the real radar hits more than this exam credits it with. The embeddings that read the contract's object are an open-weights model (Qwen3) run on a local machine, with no API calls.

Who works with whom

The same history answers another question: crossing the twelve bodies that award the most with their awarded suppliers, do patterns show up? They do, and they are sectoral and clean. It is not a ranking or an accusation: it is the picture of the split, exactly as it stands in the public records.

bodies and awarded suppliers, three years 12 bodies · 15 named firms · public data
Main named awarded suppliers by body (three years)
Body Awards Main named awarded suppliers
TRAGSA16,067dispersed split: almost all (16,053) outside the top 15
Valladolid Univ. Hospital11,916IBERDROLA 143, Becton Dickinson 143, Medtronic 139, B. Braun 115
State Central Purchasing11,632CANON 620, IBERDROLA 450, INETUM 425, ENDESA 376, CEPSA 335
Renfe Eng. & Maint.8,216Castro Perero 514, I2M 432, Faiveley 359
Paradores6,811Frutas y Verduras Massanassa 319
Insular Univ. Hospital (Canary Is.)5,780B. Braun 108, Medtronic 76, Palex 53
Cantabria Health Service4,442Becton Dickinson 181, B. Braun 150, Medtronic 148, Palex 117
ADIF (Chair)3,582dispersed split: almost all outside the top 15
Asturias Care Homes3,222INETUM 19, Medtronic 18, Telefónica Sol. 17
FREMAP3,059B. Braun 10, INETUM 8, Medtronic 8
Murcia Health Service2,809Medtronic 100, Palex 58, B. Braun 14
RTVE (procurement)2,765Telefónica Sol. 18, INETUM 11, IBERDROLA 9
How to read it: each stroke links a body with an awarded supplier, and its thickness grows with the number of awards; on hover, who works with whom lights up in amber and the exact figure appears. The groups are sectoral: hospitals and health services share the same medical-supply firms (Medtronic, B. Braun, Becton Dickinson, Palex); Renfe concentrates its spend on railway material without overlapping anyone; the State's central purchasing body splits among tech and energy majors; Paradores buys food. So the ring stays legible, only the 15 named suppliers are drawn; the bulk of each body's volume sits in an "others" bag that is not painted. That dispersion is itself a finding: TRAGSA, with 16,067 awards, spreads almost everything across a cloud of small suppliers, which is why it shows up here almost detached. Official names from the public records; no affiliation, no judgement.

Who wins it: inertia beats the model

The third delivery is the most ambitious and the one that comes out worst for an easy headline, which is why we publish it all the same. The question was whether you can anticipate which company will win a tender. The dumbest rule there is, "it goes to whoever has won it most before with that same buyer in that same sector", gets the exact name right 12.1% of the time. Our model, with the future hidden, falls short: 9.5%. It does not beat inertia, and the pre-registration already said that, if it did not, the finding would be inertia itself.

who wins: inertia rules getting the awardee right · model vs the usual supplier

12.1 / 9.5%

getting the winner's exact name right first time (inertia / the model): here the simple rule wins, and we publish it

59.2%

is the reachable ceiling: four in ten tenders are won by a company new to that body, impossible with this method. Our pre-registration estimated it at 95.6%, inflated by measuring without hiding the future

16.3%

of cases the system abstains, when there are too many candidates or too little confidence; within the 10-20% band fixed before looking

How to read it: getting it right first time is top-1; if we give three names (top-3), inertia puts the winner in 22.8% of the time and the model 20.1%; in a list of ten they tie (35.8% against 35.6%). The model beats the weak baselines, so it has learned something real, but mixing signals it sometimes overthinks and puts a bigger company ahead of the one that actually repeats. The useful reading for a company is not guessing the name, it is knowing where there is no usual supplier: that is where a newcomer really has a chance.

The caveats, from the start

What about your company?

The same reading of the history works to budget a job or a service by reading what similar ones cost instead of guessing, to know in advance which calls, offers or internal tenders will go unanswered, or to estimate how much an auction will close at before entering it.

If one of these experiments reminds you of a problem of yours, a history nobody uses, decisions that depend on a veteran's hunch, public data nobody cross-references, write to us and we'll talk it through. No fluff: we'll tell you whether it can be done.

Built with Python (streaming lxml over the official history), Parquet storage and classical machine learning (scikit-learn), with a fixed seed and pinned versions so the experiment reproduces from scratch. No language model makes the decision: it is statistics and classical models over public data.

Data and thanks

This experiment exists because Spanish public procurement is published openly. Thank you.

No affiliation with any authority or company. The names of awardees that may appear are the public ones from the procurement records themselves.

Can you use this? Yes. Our computations, the funnel and the text on this page are published under a CC BY 4.0 licence: use, share or analyse them freely, citing the source, Team Banzai (team-banzai.com). The public procurement data keeps the terms of its official source.

← All tech demos