Skip to content
team banzai

tech demo · health · public data

Connecting every medicine in Spain and showing, with the source one click away, what its official label says

We are building a search tool that relates the medicines sold in Spain. It reads the official summary of product characteristics and patient leaflet of the 16,126 marketed ones, connects them, and lets you ask in plain language to see, with the source one click away, the connections behind each answer.

The idea, in one sentence: take a heap of scattered documentation, a whole domain of it, and turn it into something you can walk through and ask questions of without ever losing sight of where each fact comes from. Here the heap is the official medicine labels, because they are public and anyone can check what we say.

What matters to a company is not this particular case, it is the pattern: this same thing, with your manuals, your contracts, your internal procedures or your technical reports, is a place where your people ask and find what is already written, connected and with the original document alongside.

You can touch it below already, on a first sample corpus that grows as the rest of the medicines come in. We set out from which source it is made, what you can touch, and the health boundary we do not cross. The interface is in English; the corpus and the quotes are in Spanish, their official language.

Let's start big. Every dot below is a medicine sold in Spain, placed near the others in its own therapeutic family (the group it belongs to by what it treats: the heart, the nervous system, infections…) and with its colour. The brighter a star, the more often the brain has named it while reading the other labels. Search for a medicine and the camera flies to its star; tap any one and it opens what its label says.

the whole brain

Real data from CIMA (AEMPS). Colour by therapeutic family (first level of the ATC code), brightness by how many times other labels name it. If the brain does not answer, the piece says so and the rest of the page stays up. On a phone it may show a reduced version.

And now one by one: the medicine and everything around it

The galaxy is the map; this is the magnifying glass. Put a medicine at the centre and you will see, pulling from the official fields, its active substance (what does the work), its lab, its family, the excipients a patient cares about (lactose, gluten) and the other brands that share its substance. And on top, in amber, what its label mentions about others: the proof opens on a click.

the brain · level 1 · exploration

Real data from CIMA (AEMPS), sample corpus. Each node links to its official label. With no connection to the brain, this piece says so and the rest of the page stays up.

What we already know from the source

Before promising anything, the count of what is there. These figures come from the official nomenclature, the public catalogue of medicines, measured on 16 July 2026.

the starting catalogue source: cima · aemps · marketed only

16,126

marketed medicines (of 20,301 authorised; non-marketed and revoked ones are filtered out)

3,365

active substances, the ingredient that does the work and that brands and generics share

1,621

marketing-authorisation holders, each with its own portfolio of products

7,351

ATC codes, the official classification that groups medicines by therapeutic use

575

declared excipients, such as lactose or gluten, that connect products to one another

2

label sections that carry the serious level: contraindications (4.3) and interactions (4.5)

The AEMPS serves the labels and leaflets already split by section as clean text, not as a flat PDF, and with a daily change log so we know what has been modified. That is why the ingestion is manageable and the information can be kept fresh.

Two ways to touch the brain, and they differ on purpose

The most important design decision: keep what can be shown with no risk apart from what touches health ground. Two levels, and the second one is handled with gloves on.

level 1 · exploration

The connections that come straight from the data, no interpretation

The one the demo opens with. It shows relations that come straight from the official catalogue, without a single word added by us: which medicines a single lab shares, which brands and generics carry the same active substance, which products share an excipient such as lactose or gluten, which have an open supply problem, what year each authorisation is from, or which therapeutic family groups which. Facts, not advice, and that is where the fun is.

level 2 · documentary

What the official label says about others, quoted verbatim

The serious one. When you ask what the label of one medicine says about another (the contraindications and interactions sections, 4.3 and 4.5), the brain neither summarises nor opines: it shows you the verbatim quote, with a link to the official document and the date it was obtained. Not a single clinical word added by us.

This shows the official documentation, it is not medical advice

This framing is not the small print at the bottom, it is part of the product, so it goes here, in plain sight. What the brain does is show you, connected and in plain language, the official AEMPS documentation. What it does not do, and will not do, is decide for you. To a question like "can I take this with that" it never answers yes or no: it shows you what each label says, with its quote and its link, and it points you to your doctor or your pharmacist.

No traffic lights. We paint no green, amber or red over any combination. The official label gives no such signal, and adding it ourselves would be inventing it.
The absence of a mention is not safety. If the document does not talk about a combination, we say so in those words: not appearing does not mean it is safe.
Each mention runs in its own direction. One medicine's label mentioning another does not mean the other's mentions it; both sides are shown separately, not merged into a false arrow.

And above all: always consult your doctor or your pharmacist. This is a tool to read what is already public more clearly, not to replace a professional.

Ask it, and it shows what the label says, quoted verbatim

This is the serious level. Type a query and the brain searches the contraindications and interactions sections (4.3 and 4.5) of the labels it has already read, and returns the official text as it is, with a link to its document and the date it was obtained. It neither summarises nor opines. And if what you ask is not in the documentation, it tells you instead of filling in. It runs on the sample corpus for now, so there are few labels with text. The query works best in Spanish, the language of the source.

try:

What does one label say about the other? The path between two medicines

The old question, "can I take this with that?", is never answered here with a yes or a no. You pick two medicines and the brain shows you, separately, what each one's label says about the other: the direct mentions, and the ones that arrive by family (one label names a whole class, and the other belongs to that class by its official code), with the chain in plain sight and the literal quote a click away. If one label does not mention the other, it says so in those words. And above all: ask your doctor or your pharmacist. The corpus and the quotes are in Spanish.

This already existed, and we say so

Building a knowledge graph from text and answering with a citation is not something we invented: it is a mature field with product behind it. Here is what was there, what we do the same and what we change.

What was there before GraphRAG (Microsoft) and LightRAG build graphs from text and answer with citations; Neo4j's builder does it as a developer tool; NotebookLM shows mind maps with a link to the page. In medicines, the AEMPS's own CIMA and Vademecum are label-by-label search tools, and there are interaction checkers that return pairwise tables.
What we do the same The engine: ingest documents, extract entities and relations, index into a graph plus vectors, and answer citing the source. None of that is ours, and we say so.
What we do differently We build it for the citizen over a public Spanish corpus that anyone can check, and we show the graph as the answer, not a summary or a pairwise table: you see the connections that support what we state.
What we add on top The answer shows the relations that support it, with the link all the way to the exact label; and when the official data does not back something up, the system says so instead of filling in. In health, that difference between staying quiet and answering too much is exactly the one that matters.

How old the data is, always in plain sight

Stale text in health is a problem, so the date and time of the last update are always visible, not hidden in the footer. It is not a "today" painted onto the page, which would lie: it comes from the last time the system actually reloaded from the official source. If that load failed or stopped, the page itself gives it away instead of hiding it, and if the data is more than a couple of days old, it says so in full.

last update sync with cima
loading…

What about your company?

The same idea works for any heap of documentation that nobody in your company quite makes use of: product manuals, tenders and contracts, internal procedures, the technical reports asleep in a folder. Instead of every query hanging on whoever has been there for years and knows where everything is, your people ask and find what is already written, connected and with the original document alongside.

If this sounds like a problem of yours, write to us and we'll talk it through. No fluff: we'll tell you whether it can be done.

The data comes from CIMA, the AEMPS's public medicines database, with a bulk download and an official API. The load runs in Python, with a full local copy so we don't call the source on every visit, and a daily sync so we don't show stale text. The direct relations are computed without any language model, from the structured data alone. Linking what one document says about another does need recognising active-substance and class names inside the text, and there we always keep the exact quoted span. The search tool will live on the same infrastructure this site already uses.

Data and acknowledgements

This demo exists because the AEMPS publishes CIMA in the open. Thank you.

Source of the information: Spanish Agency of Medicines and Medical Devices (AEMPS, www.aemps.gob.es), obtained on 16 July 2026. The official texts are shown unaltered and with a link to their document; the AEMPS neither endorses nor takes part in this demo.

Can you use this? Our texts and computations on this page are published under a CC BY 4.0 licence: use them citing the source, Team Banzai (team-banzai.com). The medicines data keeps the conditions of its official source, the AEMPS.

← All tech demos