The question nobody can answer quickly
Ask an engineering organisation what is inside the software it ships and, increasingly, it can tell you. SBOMs have made that answer routine: packages, versions, licences, known vulnerabilities.
Ask the same organisation a different question and the room goes quiet:
How was this software actually built, tested, secured, packaged and deployed?
The honest answer is usually "it's in the YAML somewhere." Delivery today is spread across GitHub Actions, GitLab CI, Jenkins and Azure Pipelines, often all at once. Pipelines call reusable workflows in other repositories, include shared templates, load Jenkins shared libraries, pull third-party actions by tag, start container images, run scanners, read secrets and deploy to environments. Multiply that by hundreds or thousands of repositories and nobody holds the whole picture.
So there is a gap. We have a structured inventory of the software components. We do not have a structured representation of the delivery system that produced them. The CI/CD Bill of Materials is our answer to that gap.
What a CI/CD BOM is
A CI/CD Bill of Materials is a structured, machine-readable description of the workflows, jobs, steps, tools, external dependencies and relationships involved in building and delivering software, together with the security-relevant facts attached to them: permissions, triggers, environments and which secrets are referenced.
An SBOM tells you what is in the software. A CI/CD BOM helps explain how that software came into existence.
The core hierarchy is simple:
Workflow → Task (job) → Step → Tool / Action / Image
with explicit edges between tasks, and from tasks to the external things they use. The point is not to produce a prettier copy of the YAML. The point is to stop treating CI/CD as a pile of configuration files and start treating it as supply-chain data.
Why structured data matters
Today, CI/CD knowledge lives in provider-specific files: .github/workflows/*.yml, .gitlab-ci.yml, Jenkinsfile, azure-pipelines.yml. Each provider describes the same handful of concepts in its own dialect. A GitHub job's needs: is a GitLab job's needs: or stage ordering, an Azure job's dependsOn:, a Jenkins stage sequence. A GitHub uses: org/repo/.github/workflows/x.yml@ref plays the same role as a GitLab include: or an Azure template:.
A tool that wants to answer questions about delivery has two options. It can learn every dialect, every time, for every question. Or the dialects can be normalised once into a common model, and every question is asked against that model.
Once the data is normalised, you can:
- search the whole CI/CD estate instead of cloning repositories one at a time;
- see which jobs depend on which, and which workflows lean on the same reusable workflow or template;
- see which external actions, images and shared libraries are in use, and whether they are pinned;
- check which kinds of control (lint, test, scan, build, deploy, release) a pipeline declares;
- spot pipelines with no scan step, or with write permissions they probably do not need;
- compare pipelines across teams on the same axes;
- feed a dependency graph for blast-radius and attack-path analysis;
- keep a snapshot of what the delivery process declared at a given commit.
The conceptual shift is one sentence: CI/CD configuration becomes queryable supply-chain data.
The data model
Before any JSON, the concepts.
Workflow. One pipeline definition: a GitHub Actions workflow file, a GitLab pipeline, a Jenkins pipeline, an Azure pipeline. It carries its provider, source path, a content hash of the source file, its trigger, its declared permissions, the environments it targets and the names of the secrets it references.
Task. A logical unit of work inside a workflow: what GitHub and GitLab call a job. Each task is classified into one or more task types (clone, clean, lint, scan, test, build, merge, deliver, deploy, release, copy, other). We classify conservatively from the job name and its commands, and we record the evidence for every classification so a reader can see why a job was called a scan. A job named deploy-preview that only builds is tagged both build and deploy; we accept that false positive because the evidence makes it explainable.
Step. One execution unit inside a task: a shell command, an action invocation, a script, a scanner call. Steps keep their order, their condition, their working directory and the names of any secrets they touch.
Resource (component). Anything the pipeline uses that lives outside the pipeline definition: an external GitHub Action, a reusable workflow, a GitLab template, a Jenkins shared library, a container image, a service container, a runner. Each resource records whether it is remote and whether it is pinned. For a GitHub Action, "pinned" means a full 40-character commit SHA; a tag like @v4 is not pinned, because tags move.
Relationships. This is where the value is. Inventory alone tells you what exists. Relationships tell you what affects what:
- Task A depends on Task B, from explicit
needs:ordependsOn:, or from stage ordering when no explicit dependency is declared. - Task uses Resource X: this job calls this action, runs in this image, delegates to this reusable workflow.
- Workflow is triggered by Event E: push, pull request, schedule, manual dispatch.
Every workflow, task, trigger and resource has a stable identifier, so relationships survive re-serialisation and can be joined across documents. That is what eventually allows graph analysis rather than file grepping.
Why CycloneDX
We could have invented a format. It would have been quicker for about a month.
The reason we did not is architectural, not a compatibility checkbox. CycloneDX is already a general-purpose model for software and system transparency: components, services, dependencies, identity via bom-ref and package URLs, metadata, and a defined extension mechanism. And it contains something most people who use it for SBOMs never touch: Formulation.
Formulation exists to describe how something was manufactured or deployed. Its vocabulary is formulas, workflows, tasks and steps, with task types, task dependencies, triggers, resource references and executed commands. That is, almost exactly, the vocabulary of CI/CD. Even the task-type enumeration in the specification (clone, lint, scan, test, build, deliver, deploy, release and so on) matches the categories we wanted to classify jobs into.
So a CI/CD BOM in our implementation is a CycloneDX 1.7 document whose substance lives in formulation[]:
- one formula per repository and revision;
- a workflow per pipeline definition;
- tasks with
taskTypesandsteps; taskDependenciesfor the job graph;triggerfor how the workflow starts;resourceReferencespointing from workflows and tasks to the external resources, which are listed as components inside the formula.
Building on an existing model buys us things a proprietary format cannot. CycloneDX tooling can open the document. Identity works the same way it does in an SBOM, so an action referenced as pkg:github/actions/checkout@<sha> in a CI/CD BOM is the same kind of identifier a security team already uses elsewhere. And the document sits naturally beside the other BOM types instead of in its own silo.
Where we extend the model
CycloneDX is not a CI/CD configuration language, and it should not become one. Several facts that matter for CI/CD security have no first-class field: which provider a workflow came from, which stage a job belongs to, the job's permissions block, its matrix, the names of the secrets it reads, the environments it targets, whether an external resource is pinned.
CycloneDX anticipates exactly this with properties: name/value pairs attached to any object, under a namespace you own. We put every CI/CD-specific fact there, under a namespace that is clearly ours, for example:
…:ci:provider,…:ci:source:path,…:ci:source:sha256on a workflow;…:ci:workflow:permissions,…:ci:workflow:environment-names,…:ci:workflow:secret-names;…:ci:task:stage,…:ci:task:matrix,…:ci:task:condition,…:ci:task:permissions,…:ci:task:type-evidence;…:ci:resource:kind,…:ci:resource:remote,…:ci:resource:pinned,…:ci:resource:runner-type.
Two rules shaped this.
First, we deliberately did not use the cdx: prefix. That namespace belongs to the specification. Squatting on it would collide the day CycloneDX defines those fields properly, and we would rather migrate from our namespace to theirs than fight over the name.
Second, everything that fits a core field goes in the core field. Jobs are tasks, not a custom "job" property. Job ordering is taskDependencies, not a custom graph. Triggers are trigger. Extensions carry only what the core model cannot.
The design principle, stated plainly: extend the standard where necessary, rather than replacing it.
A worked example
Take a small release workflow:
# .github/workflows/release.yml
on: [push]
permissions:
contents: read
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@<40-char-sha>
- run: npm ci
- run: npm run build
security:
needs: build
runs-on: ubuntu-latest
steps:
- uses: github/codeql-action/analyze@v3
- run: trivy image app:latest
deploy:
needs: [build, security]
uses: example-org/deployment/.github/workflows/deploy.yml@main
secrets:
DEPLOY_TOKEN: ${{ secrets.DEPLOY_TOKEN }}
As a tree:
Repository
└── Workflow: release.yml trigger: push permissions: contents:read
├── Task: build types: clone, build
│ ├── Step: checkout uses actions/checkout (pinned to SHA)
│ ├── Step: npm ci
│ └── Step: npm run build
├── Task: security types: scan depends on: build
│ ├── Step: CodeQL analyze uses github/codeql-action (tag, not pinned)
│ └── Step: trivy image
└── Task: deploy types: deploy depends on: build, security
└── delegates to reusable workflow
example-org/deployment/.github/workflows/deploy.yml@main (not pinned)
secret names referenced: DEPLOY_TOKEN
In the CI/CD BOM that tree stops being indentation and becomes objects and edges. A simplified excerpt:
{
"bomFormat": "CycloneDX",
"specVersion": "1.7",
"metadata": { "lifecycles": [{ "phase": "pre-build" }] },
"formulation": [{
"components": [
{ "bom-ref": "pkg:github/actions/checkout@<sha>", "type": "application",
"properties": [{ "name": "…:ci:resource:kind", "value": "action" },
{ "name": "…:ci:resource:pinned", "value": "true" }] },
{ "bom-ref": "pkg:github/example-org/deployment@main#.github/workflows/deploy.yml",
"properties": [{ "name": "…:ci:resource:kind", "value": "reusable-workflow" },
{ "name": "…:ci:resource:pinned", "value": "false" }] }
],
"workflows": [{
"bom-ref": "wf-release",
"name": "release.yml",
"taskTypes": ["clone", "scan", "build", "deploy"],
"trigger": { "type": "webhook", "event": { "description": "push" } },
"tasks": [
{ "bom-ref": "task-build", "name": "build", "taskTypes": ["clone", "build"] },
{ "bom-ref": "task-security", "name": "security", "taskTypes": ["scan"] },
{ "bom-ref": "task-deploy", "name": "deploy", "taskTypes": ["deploy"],
"resourceReferences": [{ "ref": "pkg:github/example-org/deployment@main#.github/workflows/deploy.yml" }],
"properties": [{ "name": "…:ci:task:secret-names", "value": "[\"DEPLOY_TOKEN\"]" }] }
],
"taskDependencies": [
{ "ref": "task-security", "dependsOn": ["task-build"] },
{ "ref": "task-deploy", "dependsOn": ["task-build", "task-security"] }
]
}]
}]
}
The serialisation is not the interesting part. The semantics are. From this document a machine can answer, without knowing anything about GitHub's YAML dialect: this workflow declares a scan task, and the deploy cannot start before it. The deploy delegates to a reusable workflow in another repository that is referenced by a moving branch, and it passes a named secret across that boundary. The CodeQL action is referenced by a tag, not a commit.
Two things are deliberately absent. Secret values are never read; only names are recorded. And there is no run data: this document is declared-only. It says what the pipeline definition says, not what a particular run did. That keeps one scan producing one deterministic document, which matters when you want to compare two revisions and trust that any difference is a real change.
Honest about what it cannot parse
Real pipelines are messy. Jobs reference jobs that do not exist; dependency cycles slip in through typos; some constructs are simply outside what a static parser can represent.
We made an early decision here: a dangling or cyclic job dependency degrades the document to partial instead of rejecting it. Rejecting would throw away an otherwise valid inventory of an entire pipeline because of one typo. Instead the offending edge is dropped, the reason is recorded on the workflow, and the document-level status says partial. A consumer can then decide how much to trust it. A workflow we cannot parse at all is still listed, marked unsupported with a reason code, so a gap in coverage is visible rather than silent.
Every document is also validated against the vendored CycloneDX 1.7 JSON schema before it leaves the generator, and remote schema fetching is disabled. If we cannot produce a valid CycloneDX document, we do not produce one.
What organisations can do with it
Inventory. Which pipelines exist, which providers, which external actions, images, templates and shared libraries, and how many are unpinned.
Dependency analysis. Which workflows lean on the same reusable workflow, template or shared library; which jobs gate which.
Security assurance. Does this pipeline declare a scan task at all? Which jobs run with write permissions? Which jobs receive which secrets? The task-type evidence tells you whether "scan" was inferred from trivy in a command or from a job called security, which is the difference between a control and a label.
Attack-path analysis. If a reusable workflow or a shared template is compromised, every workflow with a resourceReference to it is in the blast radius, along with the secrets those tasks hand it and the environments they deploy to.
Incident response. When a popular action or base image is compromised, the question "where do we use it, and at which version?" becomes a query over package URLs rather than an organisation-wide grep.
Historical evidence. Each document is tied to a repository revision and a content hash of every pipeline file it read, so the declared delivery process at a given commit is kept as data.
Search. Instead of opening repositories one by one, query the estate.
How it fits with other BOMs
A mature supply-chain picture has several views:
- SBOM: which software components exist?
- CBOM: which cryptographic assets exist?
- AI/ML BOM: which models and AI assets exist?
- CI/CD BOM: how was the software built, tested, secured and deployed?
Because they share one model, they share identity. The action in a CI/CD BOM and a component in an SBOM use the same kind of identifier, and both can reference the same repository at the same revision. The CI/CD BOM does not replace an SBOM. It answers a question the SBOM was never designed to answer.
The bigger idea
Software supply-chain transparency cannot stop at application dependencies.
The delivery system is part of the supply chain. Pipelines execute privileged code, hold credentials, pull third-party components at run time, perform (or skip) security checks, and push software into production. Several of the most damaging supply-chain incidents of recent years did not touch a single application dependency; they went through a build system, a CI action or a pipeline credential.
That system deserves the same structured visibility the industry now expects for dependencies. Not another dashboard over YAML, but data: normalised, identified, related, and built on an open standard so it can travel.
At KvantumCI, this is how we generate CI/CD BOMs today: our generator parses GitHub Actions, GitLab CI, Azure Pipelines and Jenkins definitions (with preliminary support for CircleCI, Tekton, AWS CodeBuild, TeamCity and Gitea) into CycloneDX formulation documents, stores them beside SBOM, CBOM, AI BOM and ML BOM, and draws each pipeline in our estate-wide map alongside the capabilities our rule engine detects. But the format does not belong to us, and that is the point.
If we want to understand the provenance of software, we cannot only describe what is inside the artifact. We also need to describe the system that built it.
See your pipelines as a CI/CD BOM
Connect a repository and KvantumCI generates a CycloneDX CI/CD BOM next to your SBOM.
Start Free Trial