The Rabbit Hole Generator gave me Macaron this week, Oracle’s open-source tool for checking software supply chain security. My first instinct was to file it next to every other dependency scanner I’ve used: point it at a package, get a list of CVEs back. That’s not what it does, and the gap between those two things turned out to be the actual point.

A vulnerability scanner answers “does this package version have a known CVE.” Macaron answers a different question: “was this artifact actually built from the source code we think it was, by a process we can verify, with evidence that isn’t just the vendor’s word for it.” Those sound similar. They aren’t. A package can have zero known CVEs and still have no evidence at all connecting the file you pip installed to the GitHub repo everyone assumes it came from.


What is Macaron used for in supply chain security?

Macaron is a policy-driven analysis tool that inspects a package’s build and release history and checks it against a set of supply chain integrity rules, most of them built around SLSA. SLSA (Supply-chain Levels for Software Artifacts) is an industry framework that grades a build pipeline from Level 0 (no guarantees) to Level 3 (a build that’s isolated, tamper-resistant, and produces cryptographically verifiable proof of what ran). Instead of asking “is this code vulnerable,” Macaron asks “can we prove this artifact came from this source, via this build, with this level of integrity,” and turns the answer into a pass/fail checklist per package.


The Architecture: a policy engine, not a single scanner

Under the hood, Macaron runs a fixed pipeline for any target you give it (a PyPI/npm/Maven package, or a raw repo): it resolves the package to its source repository, clones it, pulls whatever CI metadata and build configuration it can find, looks for a real in-toto provenance statement (a signed, standardized JSON record of what built an artifact, the same format Tejolote produces, covered in an earlier post in this series), and then runs each of roughly seventeen named checks against everything it collected. Checks have IDs like mcn_provenance_available_1 or mcn_trusted_builder_level_three_1, and some depend on others: you can’t pass “trusted builder level three” if “build as code” already failed, so failures cascade down a tree rather than sitting as flat independent flags.

The interesting design choice is that Macaron doesn’t hardcode “this package is secure” or “this package is not.” It hardcodes checks and lets you compose policy over the results. That makes it closer to a rules engine for supply chain claims than a single opinionated scanner, and it’s why the tool ships with a separate verify-policy command distinct from analyze. You run the checks once, then decide what combination of passes you actually require for your own risk tolerance.


The Reality Check: where this breaks

1. Ecosystem coverage doesn’t scale to a real polyglot monorepo

Macaron’s provenance and build-inference logic is written per ecosystem and per build tool. A shop running Python, Java, and a hand-rolled internal build wrapper in the same repo will get uneven, sometimes silent gaps in coverage, not a clean unified answer. The tool’s own docs describe ecosystem support incrementally, which is honest, but it means “run Macaron on everything” is not a one-command guarantee the way “run a CVE scanner on everything” nearly is.

2. Analysis cost grows with dependency graph size and CI rate limits

Every target means cloning a repo, walking its CI configuration, and potentially recursing into dependencies. Run this across a large dependency tree and you’re making a lot of GitHub API calls in a short window, the same rate limits that bite any tool doing live repository introspection. This is a real operational cost, not a hypothetical one: the analyze command needs a GitHub token specifically because of how much metadata it pulls per target (more on that in the lab).

3. Deterministic policy outcomes across thousands of packages is a research problem, not a shipped feature

Package metadata quality varies wildly across the open source ecosystem. A repo link that’s stale, a release tag that doesn’t match a commit, a package published from a fork instead of the canonical repo, all of these produce noisy or “unknown” results rather than a clean pass or fail. Point Macaron at a dependency tree with a few thousand entries and expect a meaningful chunk of “we can’t tell” results, not because the tool is broken, but because the ecosystem it’s inspecting wasn’t built with this kind of verification in mind.

Macaron’s whole chain starts from “here’s the package, go find its source repo.” If that link is wrong, stale, or attacker-controlled, everything downstream inherits the error. The tool is verifying a claim about provenance using data whose own provenance it doesn’t independently verify.

5. It assumes upstream CI attestations exist and are retained long enough to check

If a project never published an in-toto statement, or published one and then let the storage holding it expire, Macaron has nothing to verify against. That’s not a bug, it’s the honest limit of the approach, and (spoiler for the lab below) it’s the single most common outcome you’ll actually see running this against real-world Python packages today.

6. Introduced Vulnerability

Running Macaron in CI means giving one process broad read access to your source repositories, CI metadata, and package registry data simultaneously, all authenticated with tokens that need to reach across that whole surface. That’s a credential aggregation point: compromise the runner or its token store and an attacker inherits read access to everything Macaron was configured to see in one place, rather than having to compromise each system separately.


The 30-Minute Lab: checking a real Python package for real provenance

The generated Sandbox Challenge said to pip install macaron and point it at Django. The first command in that instruction doesn’t work, and the reason why turned into the first real finding of the lab.

1. Installation: the generated instructions point at the wrong package

I ran the sandbox challenge as written first, inside a throwaway Docker container (not on my host; whatever a pip install pulls down is code I haven’t reviewed, and there’s no reason to let it touch anything but a container I’ll throw away):

docker run --rm python:3.12-slim bash -c "
  python3 -m venv .venv && . .venv/bin/activate
  pip install --upgrade pip
  pip install macaron
"

That failed with a SyntaxError in macaron.py itself, except sqlite3.IntegrityError, e:, which is Python 2 syntax. That’s not a bug in Oracle’s tool, it’s a sign the install target is wrong entirely. PyPI’s macaron package is a dormant, unrelated SQLite ORM last released in 2012, years before Oracle’s tool existed. The name is simply squatted. The Sandbox Challenge’s generator pattern-matched “pip install ” without checking that the name actually resolves to the right project, a good reminder that even a well-written summary and Critics Corner from the generator doesn’t guarantee the one literal command it hands you is correct. Always check the tool’s own docs for the real install path before running anything it suggests.

Oracle’s real install method is a wrapper script that pulls a Docker image (ghcr.io/oracle/macaron) and runs it for you:

curl -O https://raw.githubusercontent.com/oracle/macaron/refs/tags/v0.25.0/scripts/release_scripts/run_macaron.sh
chmod +x run_macaron.sh
./run_macaron.sh --help

Two more real gotchas surfaced running just that: the script warned my Mac’s default bash (3.2 from 2007, Apple hasn’t shipped a newer one in years over GPLv3 licensing) is unsupported (it still ran fine, but it’s the kind of warning worth reading rather than ignoring); and the analyze command needs a GITHUB_TOKEN environment variable to do its repository lookups, something the generated Sandbox Challenge never mentioned at all. I already had a token from gh auth login, so I piped it in without ever putting it in a shell history or a file that outlives the run:

gh auth token | xargs -I{} env GITHUB_TOKEN={} ./run_macaron.sh analyze -purl 'pkg:pypi/[email protected]'

2. Why Django, and what the run actually does

I picked Django instead of a small package on purpose: it’s one of the most widely deployed, actively maintained Python projects that exists, with a real CI pipeline and millions of downstream users trusting its releases. If a tool built to catch supply chain gaps has nothing to say about a project this scrutinized, that would say something too. Behind the scenes, the -purl flag (Package URL, a standard way to name “this exact package from this exact ecosystem”) tells Macaron to resolve django==5.0.6 to its GitHub repo, clone the commit that version was actually tagged from, walk its GitHub Actions configuration, and run all seventeen checks against what it finds.

3. Running it, and what the output actually proves

The run finished in a few minutes and printed a summary table, then wrote a full JSON report:

 Total Checks  17
 PASSED        5
 FAILED        10
 SKIPPED       1
 DISABLED      0
 UNKNOWN       1
 JSON Report   output/reports/pypi/django/django.json

Ten failures out of seventeen sounds alarming until you look at which ten. Every provenance-related check failed: mcn_provenance_available_1, mcn_provenance_verified_1, mcn_provenance_derived_commit_1, mcn_provenance_derived_repo_1, mcn_provenance_witness_level_one_1, and mcn_trusted_builder_level_three_1 (a “trusted level 3 builder” is one that meets SLSA’s highest tier, an isolated, non-tamperable build service, roughly the bar GitHub’s own hosted Actions runners with the right configuration can meet). Their justification field in the JSON report is identical across the board: "Not Available." That single phrase is the whole story: Django, despite 89,290 GitHub stars and a real GitHub Actions pipeline, has never published a signed in-toto provenance statement for its releases. Macaron isn’t reporting that Django’s provenance is broken or forged, it’s reporting that the evidence simply doesn’t exist to check in the first place. That’s Reality Check point 5, made concrete: the single most common real-world outcome of running this tool isn’t “provenance verified” or “provenance failed verification,” it’s “there was nothing here to verify.”

Confirming that took one more look at the report, not just the summary table. The full JSON includes a provenances.is_inferred: true field, meaning Macaron didn’t find a real signed statement and instead synthesized a placeholder shape (literal "<URI>" and "<STRING>" tokens sitting where real values would go) just to show what a provenance would look like if one existed. Seeing is_inferred: true is the tell that separates “we checked a real attestation and it didn’t match” from “there was no attestation to check,” and that distinction is exactly what the flat PASSED/FAILED table above hides if you don’t open the JSON.

Two of the five passing checks are worth reading past the green checkmark too, not just accepting at face value:

  • mcn_githubactions_vulnerabilities_1 actually failed, and its justification names a specific, real finding: a PR-triggered workflow in Django’s own CI requesting pull-requests: write permission it doesn’t need. That’s a genuine, actionable finding about Django’s supply chain surface, not a hypothetical, sitting in output from a routine 30-minute lab run.
  • mcn_detect_malicious_metadata_1 passed overall, but the justification field contains fifteen named sub-signals, and two of them, closer_release_join_date and package_description_intent, came back FAIL even though the rolled-up result was PASSED. This is a good example of Reality Check point 3 in miniature: a heuristic anti-malware check on metadata alone will throw false amber signals even on an extremely well-known, entirely legitimate package, and a summary table that only shows PASSED/FAILED per check-id will never surface that nuance unless you open the underlying justification.

The takeaway from actually running this, not just reading the Summary Notion wrote: a clean provenance report on a real-world Python package is rare enough that “Not Available” should be your default expectation, not a red flag on its own. What you should key on is whether that absence changes your risk calculus for that specific dependency (is it a security-critical path, does it get pinned or auto-updated, is there a vendor SLA), not treat every “Not Available” as equivalent to “compromised.”


Where this fits in a real environment

Macaron earns its place in a CI pipeline as a periodic, async check against your direct and near-direct dependencies, feeding a dashboard or a Slack alert when a critical package’s provenance posture changes, not as a synchronous gate that blocks every build. Given Reality Check points 2 and 3 above, running it against your entire transitive dependency tree on every commit will hit GitHub API rate limits and produce more “unknown” noise than actionable signal; running it weekly against your top twenty or so direct dependencies, the ones that would actually hurt if compromised, is a more realistic starting scope. It stops being enough the moment you need deterministic pass/fail policy across a large, polyglot dependency set (point 3) or need it to catch a spoofed package-to-repo mapping on its own (point 4); for those, it’s one input into a broader review, not the whole answer.


Actionable Takeaways

  • Don’t trust a tool’s own quickstart command without checking the vendor’s real docs first. pip install macaron looks completely reasonable and installs a fifteen-year-old unrelated project instead. A five-second check against the project’s actual install page would have caught it before wasting a build.
  • Open the JSON, not just the summary table. A flat PASSED/FAILED count hid both a real finding (Django’s overbroad CI permission) behind a passing check, and a genuinely useful nuance (two amber sub-signals) behind another passing check. The interesting data was one layer deeper both times.
  • “No provenance found” is the default outcome, not the exception. Even one of the most scrutinized Python packages in existence had nothing signed to verify. Calibrate your expectations, and your alerting thresholds, around that reality instead of around a clean pass.

Provenance tooling doesn’t tell you a package is safe. It tells you exactly how much, or how little, you actually know about where it came from, which turns out to be a more useful question than most dependency scanners ever ask.