by Winch Labs

Benchmark · August 2026 · rebaselined September 6, 2026

The OSS Gauntlet: 40 real repos vs. the gate

A review gate earns trust by what it doesn't say. So we ran Deltz's deterministic gate — offline, no AI, no API keys — over 40 real open-source Terraform, Terragrunt and Crossplane repositories and published everything, including the three false-positive classes we found in our own rules and fixed along the way.

0parse failures / 40 repos
56findings across the corpus
28/40repos got silence
239conventions learned

Everything below is the paper trail behind those four numbers: the corpus we ran, how we got here, what survived, and what we're still honest about.

The corpus

Thirteen terraform-aws-modules repos (vpc, eks, rds, s3, iam, security-group, alb, lambda, ecs, cloudfront, atlantis, ec2-instance, route53), five terraform-google-modules, two Azure official modules plus Azure AKS and Claranet, four Cloud Posse modules, four Terragrunt trees, futurice/terraform-examples, aws-ia/eks-blueprints, awslabs/data-on-eks, philips-labs/terraform-aws-github-runner, 18F/identity-terraform, a GCP Cloud Run module, hashicorp/learn-terraform-import, and three Upbound Crossplane configurations — forty in all, each pinned to an exact SHA in repos.lock.

Headline results

Parse failures: 0 of 40. The HCL and Crossplane parsers survived every repo.
Findings: 56 across the corpus — 1.4 per repo, each spot-checked as a true statement about the code. The corpus grew from 30 repos to 40 and the count did not move: it was 54 before and 54 after, because ten more real repositories produced zero new findings. It has moved twice since, for reasons that have nothing to do with corpus size: 54 → 34 on August 30, when Crossplane composites left the claim gate, and 34 → 56 on September 6, when the classifier was fixed and they came back — attributed below, finding by finding.
28 of 40 repos got silence. A clean repo gets no comment; that's the product promise, measured.
Conventions learned: 239 across the corpus (from 61 when we started — 12 deterministic deriver families, now derived identically by the CLI scan and the server), every one carrying file:line evidence.
Runtime: about a second per repo (max gate time 1.03 s on the 4 October reproduction, 0.98 s on the pinned 6 September run — both 18F/identity-terraform). Convention derivation is separate and faster: 0.29 s at worst. A timing observation on one laptop, not a performance claim.

The part vendors usually skip: our false positives

The first run produced 622 findings. That number was wrong, and the reasons are instructive.

Class 1 — convention.unpinned_module_source, 538 → 9. The rule looked for a version inside the source URL, so registry modules correctly pinned via a version = "~> 5.0" attribute were flagged as unpinned. After the fix (local paths, ?ref= git pins, and version attributes all count as pinned), the 9 survivors are genuinely unpinned — we checked each one.

Class 2 — crossplane.missing_label, 54 → 15. In repos with no XRDs, anything carrying spec.parameters was classified as a Crossplane claim — which swept in Gatekeeper policy files and demanded cost-center labels on CIS constraints. The classifier now excludes well-known non-Crossplane API domains (gatekeeper.sh, kyverno.io, argoproj.io, karpenter.sh, …). The 15 survivors are real label gaps in actual Upbound claims.

Class 3 — Upbound package manifests, 12 → 0. Our provider detection mapped any *.upbound.io API group to a cloud provider — including meta.dev.upbound.io, which is Upbound's package metadata. So each Upbound repo's own upbound.yaml was classified as a managed resource and told to add cost-center labels to its manifest. All meta.* groups are metadata now. In the same pass, claims under examples/ started being classified correctly and their (true) label findings joined the count — so the total stayed at 54 while its composition got strictly more honest.

All three fixes shipped with regression tests. 622 → 54, on the original 30 repos. That is the 91% reduction we quote, and it describes this tuning step only: the corpus then grew to 40 without moving the count, and the later 54 → 34 → 56 is a different kind of change — attributed in its own section, never folded into this one.

The one we have not fixed

lib.aws.kms_key_rotation fires on asymmetric KMS keys, which cannot rotate. It accounts for 7 of 30 findings from the catalog profile. AWS does not support automatic rotation for asymmetric keys, so the rule is asking for something the provider refuses — the finding is true about the attribute and useless to the reader, which is our own definition of a false positive.

It is not fixed, and it is on this page for that reason. The check language has no conditional that can express "unless the key spec is asymmetric", so fixing it properly means extending the rule DSL rather than special-casing one rule. We would rather publish an open false positive than let this page imply the list is closed — a page about our false positives that omits the one we know about would disprove its own argument.

Also open, and a genuine judgement call rather than a bug: secops.s3_unencrypted fires on buckets that declare no BucketEncryption. Since 2023 S3 encrypts new objects by default, so the rule is asking for an explicit declaration of something already true. We think stating it explicitly is right — it survives a bucket policy change — but we understand the argument that it is noise, and we have not settled it.

54 → 34 → 56: what moved on August 30 and September 6, and why neither step is noise

The 54 stood from August 18 to August 30. Two changes to the gate on August 30 took it to 34, and neither measured this corpus when it landed; we attributed the step on September 6 by rebuilding one binary per candidate commit and re-gating the four repos whose counts changed. The arithmetic is 54 − 23 + 3 = 34. The attribution surfaced a product decision, the decision was made the same day, and the count moved again: 34 + 22 = 56. Both steps are below, in the order they happened.

−23: Crossplane composite resources stopped being gated as claims. A document whose kind an XRD defines was recognised as the composite and ran through no claim rule. All three Upbound repos in the corpus are v2-style — their XRDs declare no claimNames — so their hand-authored examples/*-xr.yaml files are exactly that document, and every finding on them left at once: 22 Warn-level convention findings (missing labels, a namespace outside the allowlist, naming, a storage floor) that are still true statements about those files, and one Critical plaintext_secret that was a false positive — the trigger was a committed Secret manifest in the same file, attributed to the wrong document. That was a coverage gap, and we published it as one. In a Crossplane v2 repo the composite is the hand-authored object, and for a week the gate said nothing about it. None of the 22 was a security finding; it was a product decision still to make — not a fix, and not noise. Two crossplane.namespace conventions left with the claims they were derived from: 239 → 237.

+3: a vendored CloudFormation template became visible. 18F/identity-terraform carries an AWS Quick Start template under a .template extension that no parser read before the CloudFormation frontend arrived. Three AWS::S3::Bucket resources declare no BucketEncryption, and secops.s3_unencrypted reports each, with the same SSE-S3 caveat as above. 18F leaves the silent set; the two Upbound repos join it: 29 → 30.

Both binaries were rebuilt from source and reproduce 54 and 34 exactly on the pinned checkouts, and the 54-era baseline is kept beside the later ones. The finding-by-finding table is in the corpus repository's REPORT.md, section "54 → 34: composites leave the claim gate, CloudFormation arrives".

34 → 56: the composites come back as the objects they are (September 6)

The decision was made the day the gap was attributed and shipped in garboard v0.1.65: the XRD decides what its composite resource is. A family that declares claimNames keeps the August 30 behaviour — its composite is what a claim resolves to, its labels are crossplane.io/claim-* backrefs, and it runs through no claim rule; so does any composite carrying such a backref or a spec.claimRef, whatever its XRD looks like, because a backref only exists on a derived object. The render-stream case the August 30 change fixed is unchanged, and its tests still pass. A family that declares no claimNames — a Crossplane v2 XRD, or a v1 XRD that never offered claims — has no claim at all: its composite is the object a platform user writes by hand, in a namespace, with labels and spec.parameters, and it now runs the claim rules as the claim-shaped unit it is. Findings say "Composite resource" where they used to say "Claim"; rule ids and wording are otherwise byte-identical.

+22, and the 23rd stays gone. configuration-aws-database 5 → 17, configuration-aws-network 0 → 5, platform-ref-aws 0 → 5: every one of the 22 is the same Warn-level statement about the same line the August 30 table called true — three missing labels, a namespace outside the allowlist and a name without the -<namespace> suffix on each of the four hand-authored composites, plus a storage floor on the two database examples. The plaintext_secret false positive on postgres-xr.yaml:12 does not return: the document-scoped content scan from the same August 30 commit keeps it gone now that the file is read again, and a unit test pins that fixture. The two crossplane.namespace conventions return with the objects they are derived from: 237 → 239. Silent repos 30 → 28. Diffed against the 34-era baseline, the run reports exactly those rows and nothing else — every other repo's findings, per-rule counts and conventions are identical, and all three negative-control repos stay at zero Crossplane findings. The 34-era baseline is kept beside the 54-era one and the current run.

This is re-admitted coverage, not noise added: the gate now says about those three trees exactly what it said before August 30, minus the one finding that was wrong. Net against the 54 that stood before either step, +2 — three true positives on 18F's vendored CloudFormation template, one false positive fixed. The table is in REPORT.md, section "34 → 56: composites come back as the objects they are".

CloudFormation, reported separately

The CloudFormation frontend was added after the run above, and its three template repositories are a different corpus that we do not average into the numbers on this page. Combining them would move the headline signal rate while describing a different kind of repository.

Three repos — AWS's own reference template set, widdix/aws-cf-templates and one more — pinned like the rest: 270 → 242 findings after two false-positive fixes ({{resolve:secretsmanager:…}} read as a hardcoded secret, which was wrong on all 8 hits, all of them in AWS's own templates; and Aurora cluster members read as standalone databases). 0 of those 3 repos are silent — expected of template libraries whose whole purpose is to demonstrate every service, and precisely why they are reported apart from the 28-of-40 figure above.

What survived (spot-checked examples)

A literal database password in terraform-aws-rds's examples/complete-mssql/main.tf:142. Unencrypted S3 in eks-blueprints' kubecost pattern. A 0.0.0.0/0 ingress in a public-ingress example. An unpinned registry module in futurice's Camunda stack. Findings without file:line evidence get dropped by the ranker — these all carry theirs.

What we're still honest about

Convention depth is capped by the deriver list: every learned convention is a deterministic parse, so a repo's implicit standards beyond our 12 families go unlearned for now. Two gaps we can name precisely: Crossplane derivation is thin — 3 conventions across 3 repos, one each, against 27 findings on those same repos (17, 5 and 5 on the three Upbound configurations, now that their hand-authored composites are gated as the claim-shaped objects they are), so the gate has far more to say about those trees than it has learned from them — and two deriver families fire on nothing in this corpus at all.

The list keeps growing, and this corpus now gates our releases: it is pinned to exact commit SHAs, runs weekly in CI, and the numbers on this page fail a build if they slip.

Run it on your repo

The free plan does exactly what this benchmark did — deterministic gate, your own repo, silence if it's clean: install the GitHub App. Or book a walkthrough and we'll run it live on a repo you pick.