incubator: goqu

Measurement methodology

goqu rates Go packages against one baseline, the waterlevel: the Go standard library. Every score answers the same question, how does this code measure against the code Go itself ships.

The pipeline

Packages arrive from the local module cache, from website submissions, and from released version sweeps. A submission resolves through the Go module proxy at its newest version; the cache sweep takes the newest cached version of every module; the release sweep takes the newest patch of every major.minor line and the ten newest releases in full. Any version with a dash suffix (release candidates, platform tags) is a prerelease under semver and is not scanned; a module with no tags at all is scanned at the pseudo version the proxy resolves for it.

Every version is read once. splint parses the tree into a document holding every package, symbol, doc comment and per function cyclomatic complexity, and splint sizes renders every file with its byte size off that document. Both payloads are stored per version, so ratings are reproducible and history accumulates.

A release that fails to extract in place is retried from a writable copy with a minimal go.mod; that sacrifices dependency resolution but keeps the syntactic extraction the checks need. A failure that survives the retry is recorded and shown on the package page.

The waterlevel

The baseline is the standard library source of the toolchain that ran the scan (go1.27.1), filtered to the library alone: no cmd/, no testdata, no vendored code. It is rebuilt when that toolchain changes, so the version cited here is the one the current baseline was measured with. The current waterlevel values:

MeasureValue
median file size3987 B
median package size22072 B
average cyclomatic complexity4.20
95th percentile cyclomatic complexity15
godoc coverage of exported symbols74%
test to symbol pairing78%
Go files measured2891

An exported symbol means an exported type, const, var or function, where a method only counts when its receiver type is exported too: an exported method name on an unexported type is interface plumbing, not API. The size figures count every Go file of the tree, test files included. All rating arithmetic is integer, in percent and permille units.

The checks

file_size_curve

Files are bucketed by size, doubling from under 1 KB to under 256 KB, with one bucket past the end. The cumulative share of files under each threshold is compared bucket by bucket against the stdlib curve, and the signed gaps are averaged over all ten buckets, in permille. Small files are the desired shape: smaller files, smaller context. Matching the curve scores 100; every 3 permille of gap moves the score one point, up for a larger share of small files, down for a smaller one. The last cumulative share is 1000 permille on both sides by construction, which slightly dampens the average; the slope absorbs it.

median_file_size

The package's median file size against the stdlib waterlevel of 3987 B, scored as 100 * waterlevel / median. The median is the upper median, the value at index count/2 of the sorted sizes. A smaller median than stdlib rises above 100, a larger one sinks below.

median_package_size

File sizes summed per package directory, the upper median across directories, against the stdlib waterlevel of 22072 B, scored as 100 * waterlevel / median. Small, focused packages rise above 100; sprawling ones sink below.

godoc_coverage

The share of exported symbols carrying a doc comment, scored relative to the stdlib waterlevel of 74%: 100 * coverage / 74. Matching stdlib's discipline is 100; documenting more of the API than stdlib does rises above it.

test_ratio

Test functions per exported symbol. A test function is a Test* function in a test package or in a _test.go file; benchmarks, examples and fuzz functions do not count. The ideal is one test per symbol; the stdlib pairing sits at 78%. A pairing of 80-100% scores the full 100 and earns no bonus. Overtesting is penalized proportionally: the score falls one point per percent over 100, reaching 0 at 200%, two tests per symbol. Undertesting is penalized the same way, falling proportionally from 100 at 80% pairing to 0 at no tests.

complexity

The average cyclomatic complexity of every function outside test files, exported and unexported alike, scored against the stdlib waterlevel of 4.20: 100 * waterlevel / average. Simpler functions than stdlib's rise above 100.

complexity_p95

The cyclomatic complexity at the 95th percentile of the same functions, the heavy tail, against the stdlib waterlevel of 15: 100 * waterlevel / p95. The average check rewards overall simplicity; this one penalizes hiding a few deeply branched functions behind a good average.

Scores, bonuses, the overall rating

Six checks score relative to the waterlevel and may beat it by up to 10 points: the score is capped at 110, and the quality table shows the overhead in a separate bonus column next to the base score. test_ratio is banded instead, and its ideal band is worth a flat 100. The overall rating is the integer average of the seven checks, so a package that beats stdlib in every dimension rates up to 108/100. The standard library scores 100 on the six relative checks by definition, 98 on test_ratio, and an overall of 99.

Failures never hide: a package whose newest release cannot be scanned keeps its last successful rating, with the failure reason shown on its page. Ratings recompute whenever a new version is scanned, and the history chart keeps one rated point per release line, plus the ten newest releases in full.