Second of two articles about software bills of materials. The inventory is in place — what it can answer, and what it cannot: vulnerabilities, licenses, and provenance, first by hand and then through APIs.

Where Part 1 Left Off

At the end of the first part, a document is on the table with three properties that cannot be taken for granted.

It is complete enough: 218 components from two generators, because a Maven tool does not see the npm side of a Vaadin application, and vice versa. It is reproducible: two builds from a fresh checkout deliver the same numbers, ever since a lockfile went under version control. And it sits next to the artifact, not in the build directory — in both delivery formats, for just under 0.3 percent of the transfer size.

If you have not read Part 1, you lose nothing here. It is enough to know: the document describes what was shipped, and it describes it more precisely than a document that knows only one half.

What Can Now Be Answered

A bill of materials is a list. On its own it answers no question — it makes questions answerable that previously could not even be asked:

  • Does my delivery contain a component with a known vulnerability?
  • Does one of the included licenses conflict with the planned distribution?
  • Where does the code I ship come from?

All three come up not on the day of the build but weeks later — and then under time pressure.

What the Document Alone Does Not Deliver

The list names names and versions. Whether an advisory exists for one of those versions is not in it, and it cannot be in it: advisories come into being later than the document. A bill of materials from May knows nothing about a vulnerability from August.

From this follows the structure of this article. Chapter 2 answers the first question without any tooling — just the document and a public database. Only Chapter 3 shows what a platform delivers beyond that, and only then can you judge whether the difference is worth it.

The other way around, it would be a product demo.

Am I Affected? — Without a Service

The question can be answered with on-board means. Two things are needed: the document and a public vulnerability database that understands package identifiers.

The identifier is the key. Every component in the document carries a purl — a string that unambiguously combines ecosystem, name, and version:

code
pkg:maven/org.jsoup/jsoup@1.23.1?type=jar
pkg:npm/%40vaadin/react-components@25.2.7

This makes the query a translation task, not analysis work. The identifiers are read from the document and held against the database:

python
purls = [c["purl"] for c in sbom["components"] if c.get("purl")]
abfrage = {"queries": [{"package": {"purl": p}} for p in purls]}
# POST https://api.osv.dev/v1/querybatch

For the inventory examined here: 217 identifiers, not a single advisory. The result matches what the platform says later — so the basic answer costs nothing but a few lines.

The Cross-Check

A result with no hits proves little as long as you do not know whether the query would deliver hits at all. Until recently, the same application carried an older version of one library:

code
1.22.1:  1 Meldung   GHSA-pmhh-3w7g-xqp8  MODERATE  behoben in 1.23.1
                     "Cleaner may expose markup with custom raw-text elements"
1.23.1:  0 Meldungen

The query works, the inventory is clean, and the difference between the two versions is a single digit. This is exactly why the bill of materials was collected.

Where It Ends

What is missing here is not a detail — it is everything that turns information into a decision:

  • No urgency rating. That an advisory exists says nothing about how likely it is to be exploited.
  • No indication of known exploitation. Whether a vulnerability is actually being attacked is recorded in a different catalog.
  • No history. The query describes today. Whether the same component was already affected three releases ago remains open.
  • No comparison across multiple deliveries. Anyone operating twenty applications repeats the same script twenty times.

The last point is where the manual approach tips over. For a single application, the query is a quick exercise. For an inventory that grows over years, it is a task with its own maintenance cost — and that is where the question begins of what a tool can take off your hands.

Vulnerabilities with a Platform

The same inventory, this time uploaded instead of queried. What a platform delivers on top is best shown with the version that still carried the finding — both sit there side by side.

code
GHSA-pmhh-3w7g-xqp8   jsoup 1.22.1   MEDIUM   CVSS 4.7
CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:C/C:L/I:L/A:N
EPSS 0,0019 (Perzentil 8,8)   bekannt ausgenutzt: nein   behoben in 1.23.1
"Cleaner may expose markup with custom raw-text elements"

Four of these data points were not available in Chapter 2: the vector, the probability score, the indication of known exploitation, and the fixing version. The third value in particular changes the urgency — an advisory with a percentile of 8.8 is something different from one that is under active attack.

The Before-and-After Comparison

Version with legacy componentsVersion with lockfile
Components248218
Advisories1 (MEDIUM)0

The second row is not the platform’s achievement but maintenance’s: the library was lifted to the fixing version. The platform’s part in it is a different one — it keeps both states side by side, so the path from one row to the other remains traceable. That is exactly what a single query does not deliver.

The Honest Finding

The finding sat in the Maven half of the document. The npm half, which Part 1 added with some effort, was and is free of advisories.

That is inconvenient because it contradicts the obvious narrative: the merge from Part 1 uncovered no finding here that a Maven-only SBOM would have hidden. Anyone advocating the merge should not claim it brought a lucky find to light.

The case for merging rests on coverage, not on a finding. A half you do not check is not clean just because you have not checked it.

That the npm side carries no advisory today is a statement about today — and it is only possible because that side is in the document at all.

What the Count Reveals

A detail on the side that calls for caution: depending on which access path is queried, the same platform reports one finding for the same inventory in one place and two medium and one low in another. And through one interface the version appears as 1.22.1, through the other as 1.22.2 — the delivery directory contains 1.22.2.

Both have been reported and do not fundamentally call the answers into question. They are, however, an argument for not adopting a number unchecked just because a tool delivers it — the same skepticism Part 1 applied to the self-computed document.

Supply Chain Data for Agents

Next to the regular API stands a second one that is likely the more interesting for this readership: an MCP server. The Model Context Protocol is the way a language model calls tools — and supply chain data is an obvious subject for it, because the questions asked of it are almost always phrased in natural language.

Access is stateless, authentication the same as for the regular API. 41 tools are offered:

code
get_bom_stats            get_bom_vulnerabilities     get_org_vulnerabilities
get_bom_components       get_bom_dependencies        check_component_vulnerabilities
get_license_problems_count  get_org_restricted_licenses  get_ai_license_analysis
get_provenance_results   get_org_provenance_countries   get_org_supply_risk
get_quality_gate_detail  list_audit_logs             whoami

This turns “Are we affected?” into a question that can be asked by someone who knows neither purl nor query syntax.

A Finding on the Side

Comparing the two access paths, something stands out that is unlikely to be intended: the MCP server hands out information that the regular API refuses to the same key. Who is this key, which organizations exist, what is the organization-wide vulnerability picture — answered over one protocol, acknowledged with a rejection over the other.

This is not a security hole; the data belongs to the owner of the key. But it shows that two access paths to the same data have two permission checks — and those can drift apart.

A Precaution That Deserves Attention

Third-party content comes back marked:

json
"description": "<untrusted>jsoup: Cleaner may expose markup
                with custom raw-text elements</untrusted>",
"component":   "<untrusted>jsoup</untrusted>"

The reason is immediately convincing once you have seen it. The description of a vulnerability is text someone else wrote. Presented to a language model, it is input from a third party’s hand — and could contain instructions instead of a description.

Anyone feeding supply chain data to a language model is feeding it text from the supply chain. The marker says: this is content, not an instruction.

For an article about supply chain security, that is an elegant twist. The chain whose trustworthiness is at stake reaches all the way into the description texts used to examine it.

Licenses: What the Traffic Light Reports and What It Overlooks

Twelve different licenses are spread across 218 components. The distribution is unremarkable — and that is exactly why the close look pays off.

LicenseComponents
Apache-2.0159
EPL-2.040
MIT16
EUPL-1.214
BSD-3-Clause10
GPL-2.0 with Classpath Exception7
others (LGPL-2.1, CDDL, Bouncy Castle, Apache-1.0 …)1 each

The automated check reports one restricted license and no conflict.

License distribution across 218 components, EPL-2.0 flagged as restricted

Figure 1: Twelve licenses, one flag — and one that carried no identifier.

What Gets Reported

Flagged is EPL-2.0 with 40 components — because of the commercial contributor’s indemnification obligation on commercial distribution. Affected are, of all things, the servlet container and the storage layer — exactly the building blocks an application absorbs when it ships the server itself.

That is a useful report. It calls not for panic but for a decision: anyone who distributes under these terms should know the obligation.

What Was Not Reported

More interesting is the case the same check overlooked in the same application — visible only because both versions sit side by side. The older one contained a charting library under a commercial license. Its license entry did not name an identifier but an address:

code
"license": { "url": "https://www.highcharts.com/license" }

A check that holds identifiers against a list finds nothing here it could compare — and consequently reports nothing.

The check reports the license it understands and passes over the one it does not. The second is the commercially riskier one.

The Uncomfortable Punch Line

That this component is no longer in the document today is no achievement of the license check. It disappeared because a lockfile was introduced and the package tree has since contained only what the project declares — the library was a leftover nobody had noticed.

A traffic light that knows only identifiers would have carried it along indefinitely without once turning yellow. For practice, a simple rule follows: components without an SPDX identifier are not an edge case — they are the spot where you have to look yourself. There are usually few of them; here it was exactly one.

Where Does the Code Come From?

The third question is the one that remains practically unanswerable without tooling — and the one that has been asked regularly since the European cyber resilience legislation.

It is not “who published the library,” but: who contributed to it, and from where? The answer is not in the SBOM. It is derived by evaluating the contribution histories of the source projects.

The Result for This Inventory

code
evaluable: 34 of 218 components

United States       2,513      United Kingdom          367
Germany               667      France                  309
Unknown               503      Japan                   303
China                 502      Canada                  252
Russia                423      India                   235
Contributors by country, eleven countries with sanctions relevance flagged

Figure 2: Eleven of the 89 countries appear on a list — the two largest of them in third and fourth place.

The numbers are contributors, not components. They show a picture most projects are likely to share: a core shaped by the English-speaking world, a substantial European share — and, with China and Russia, two countries that raise questions of their own in a risk assessment.

Why This Is an Indication, Not Evidence

Three caveats belong to every one of these numbers, and anyone who omits them is using the evaluation wrong:

Coverage is low. 34 of 218 components could be evaluated — for the rest, accessible histories are missing. The table says nothing about more than four fifths of the inventory.

“Unknown” is the third-largest group. 503 contributors without an attribution rank ahead of everything that follows. Every statement about shares has this blur built in.

The attribution itself is an estimate. It rests on time zones, name forms, and address fragments — attributes a person can change, and which, across a distributed project, say nothing about who has the authority to give instructions.

A contributor with an address in some country is no proof that someone there exerts influence on the software. It is an indication of where taking a closer look might pay off.

What It Is Still Good For

For a discussion about dependence on individual regions, such a distribution is the only available entry point — by hand, it cannot be compiled for 218 components. It should be used for what it is: a map, not a verdict.

And as an occasion to ask the other question no tool answers: how many of these components would you miss if they were no longer maintained as of tomorrow?

Cryptography and Provenance: Two Evaluations with Reservations

Two evaluations go beyond what the chapters so far have shown. Both answer questions that remain unanswerable without tooling — and in both cases the answer is weaker than its numbers suggest.

A side note on availability, because it contains a more general lesson. Both evaluations were initially unavailable — the API responded with the explanation that these endpoints are fundamentally closed to access keys. That sounded like a product decision no key could remedy, and it was planned for accordingly.

It was not true. In fact, the key was missing two permissions. With an extended key, the same endpoints deliver data without complaint, over both access paths.

An error message that says “not possible, period” where “you are missing a permission” would be correct costs the recipient days — they stop looking.

In practice this means: when a rejection sounds like a product boundary, asking back pays off before you strike the evaluation from the plan.

Cryptography: A Finding That Reports Uncertainty

code
Risikowert       12 / 100        Abdeckung   100 %
Krypto-Bausteine  1              Funde        1 (mittel)

The one building block is the BouncyCastle library in version 1.84. It was recognized not because anyone had read its code, but by its name — the rule amounts to “known cryptography library in the inventory.”

What matters is what the finding says. It does not say “weak cryptography is used here.” It says:

Check whether this cryptographic capability is actually used at runtime, and record the mapping where it is.

The rationale names three drivers, and the third is the real one: actual usage at runtime is not declared. Consistently, the finding’s category is not “vulnerability” but coverage, its risk type “uncertainty about cryptographic capabilities.”

This reaches the same limit Part 1 already marked, just one level deeper: the bill of materials knows the library is there. It does not know which algorithms run. Anyone who wants to know that needs a dedicated cryptography bill of materials with a runtime reference — and no tool writes that on its own.

Provenance: A Number That Needs Explaining

code
Score             43 / 100  (medium)
Components with origin    22 of 38 evaluated
Contributors with origin 427 of 3,601
Countries                 11 of 89 on a list
Lists                     ICTS · ITAR · OFAC

Broken down, the distribution is very uneven:

CountryContributorsComponentsLists
China21417ICTS, ITAR
Russia17921ICTS, ITAR
Iran1511ICTS, ITAR, OFAC
Lebanon, Hong Kong4 each5 – 6ITAR
Ethiopia, Belarus3 each5 – 8ITAR
Myanmar, Libya, Venezuela, Haiti1 – 22 – 5ITAR

Iran leads the rating because the country appears on all three lists — with fifteen contributors out of a total of 3,601.

Why 22 of 38 Does Not Mean What It Seems to Mean

A component counts as affected as soon as a single contributor from a listed country has worked on it. For projects with hundreds of contributors over two decades, that point is reached quickly. That more than half of the evaluable components are flagged this way therefore says something about the nature of open-source development first — and about risk only second.

On top of that come the caveats from Chapter 6, here with greater weight: 38 of 218 components are evaluable at all, the mean attribution confidence is 0.68, and the attribution rests on attributes a person can change.

Sanctions lists govern business relationships, not the origin of contributions. A contribution from a listed country is not a legal violation, and this evaluation is not a legal review.

What it delivers is a map: anyone who has to write a risk assessment for a government agency or a major customer has an entry point here that they could not compile by hand for 218 components. Anyone who reads it as a verdict makes wrong decisions about people.

Where a Platform Pays Off — and Where It Does Not

At this point, the question Chapter 2 raised can be answered. The basic answer — am I affected? — costs a few lines and no service. What cannot be built in-house is this: a maintained mapping from packages to contribution histories, a continuously updated assessment basis, and the running record across many deliveries.

Anyone who wants to know once whether an advisory affects them does not need this. Anyone who has to prove for years what they shipped will either buy it or build it themselves — and the in-house build is more expensive than it looks.

Obligations, Formats, Recommendation

Up to this point, the bill of materials has been a good idea. For a growing share of manufacturers it becomes a requirement — the European legal act on cyber resilience demands, among other things, that for products with digital elements the manufacturer knows and documents the included components.

Two points on this, without any claim to be legal advice: the reporting obligations for actively exploited vulnerabilities have applied since September 2026; the remaining requirements follow at the end of 2027. And it is not only those who sell software who are affected — anyone who makes it available in the course of a commercial activity falls under it as well.

Anyone who starts now has the advantage of being able to do it properly without deadline pressure.

The Four Formats, Four Axes

FormatCompletenessVerifiabilityQueryabilityEffort
WAR on an installed serverlow — the server is missingpartialsamelow
Thin Distributionhighcompletesamelow
Fat JARhighnonesamelow
jlink distributionhigh, without the runtimelike Thinsamemedium

The third column is the same everywhere, and that is not a flaw in the table: queryability does not depend on the packaging. It depends on whether the document is complete and whether anyone asks it questions. Both are independent of whether an archive or a directory tree ships in the end.

Anyone who has the choice and needs auditability takes the Thin Distribution — it is the only format in which every line of the document can be held against a file in the release.

The Objection That Belongs Here

The article series these measurements come from pursues one direction: with every step, more responsibility moves from the server into the release. Less installed software, fewer assumptions about the environment, less dependence on others.

Handing an SBOM to a service runs counter to this direction. The objection is legitimate and deserves an answer rather than a shrug.

The Resolution

It lies in the order in which this article is built.

The document lives in the release. It is usable without any service — Chapter 2 showed that the basic answer costs a few lines and delivers the same result.

The platform answers the questions that go beyond a single release: comparisons across versions, ratings by exploitability, license checking across the entire inventory, provenance estimation. Anyone who does not need them loses nothing.

The inventory belongs to whoever ships. A service is an ingredient, not a prerequisite. Whoever turns that around has bought a new dependency in order to document an old one.

What Remains

Three sentences from two articles:

The SBOM describes the dependency graph, not the delivery.

An SBOM covers what its generator can see.

A bill of materials that cannot be reproduced proves nothing.

All three are inconvenient, and all three were measured on a single application you can rebuild yourself in half an hour. That is the real recommendation: do not believe the bill of materials — hold it against your own artifact once. The surprises come on their own.