If your CodeQL wrapper reads plain-text stderr, check its severity matching before accepting a CLI upgrade. CodeQL 2.27.2 adds explicit error and warning prefixes to more messages; that does not make every stderr line a failed scan. Keep process exit status, diagnostic text and analysis results as separate evidence.
The useful question is not “Did the console look different?” It is “Did our wrapper reach the same intended decision from the changed output?” This guide provides a small parser fixture and a comparison record. It is a proposed review procedure, not a report of running CodeQL against a production repository.

What changed, and what did not
The CodeQL 2.27.2 release notes, dated October 7, 2026, say error and warning messages on standard error now carry ERROR: or WARNING: labels. Previously, only diagnostics had such labels. The notes explicitly separate this change from structured output: log files, SARIF and stored diagnostics keep severity without the added prefix. GitHub's October 9 announcement also describes the update.
That is a reason to inspect text consumers, not a promise that every CodeQL output format is permanently stable. Nor does it mean every warning is a newly discovered vulnerability. A message about running a tool and an analysis finding are different things.
Find the consumer before changing the parser
Search your own wrapper and workflow for the command it launches, the stderr capture, and the condition that marks a run failed. Follow the data all the way to the displayed status. Do not change security gates merely to make a newly red job green.
- Human-only console: check readability; a parser migration may not be needed.
- Exact text matching: test whether an added leading label defeats a match.
- Any stderr means failure: check the policy deliberately. A progress or warning line is not interchangeable with process failure.
- Structured result consumer: identify the actual file format and schema used. Avoid rewriting it solely because a console prefix changed.

Use a tiny fixture to expose a brittle assumption
The following Python example uses invented messages. It does not invoke CodeQL, establish its exit codes, or implement a production parser. It demonstrates one narrow bug: a prefix can break a match anchored at the start of a message. Save it as a temporary local file and run it with Python 3.
samples = [
("Sample extension missing", "unknown"),
("WARNING: Sample extension missing", "warning"),
("ERROR: Sample input invalid", "error"),
("Sample operation finished", "unknown"),
("message mentions ERROR: later", "unknown"),
]
def label(line):
if line.startswith("WARNING:"):
return "warning"
if line.startswith("ERROR:"):
return "error"
return "unknown"
old_match = lambda line: line.startswith("Sample extension missing")
assert old_match(samples[0][0])
assert not old_match(samples[1][0])
for line, expected in samples:
assert label(line) == expected
print("5 synthetic label cases passed; no scan was run")
The expected output is exactly the printed sentence. “Unknown” is intentional: an unlabelled line is not automatically success. The example also refuses to promote a label found in the middle of a message. Real output may contain formatting, continuation lines or messages from other processes, so test actual captured output before designing normalization. Do not blindly remove every occurrence of ERROR: from a log.

Keep three observations in the review record
| Evidence | Record | Do not substitute |
|---|---|---|
| Process outcome | The invoked command's exit status and wrapper handling | The mere presence of stderr |
| Diagnostic text | Relevant redacted lines and parser classification | A count of vulnerability findings |
| Analysis result | Expected output artifact, format and consumer result | A reassuring console sentence |
Capture the status of the process you mean to evaluate, not just the final command in a shell pipeline. Keep source revision, CLI version, query selection, configuration and command arguments with the observation. Remove credentials and private source content before sharing evidence. When comparing versions, preserve the inputs and record every intentional difference; a changed analysis result needs investigation rather than automatic blame on message formatting.

Choose an output route for the job
For a single query, GitHub's query run documentation recommends writing BQRS output for further processing and decoding it into a machine-friendly representation. It points to database analyze for alert-oriented SARIF output. These routes serve different purposes; choose the command and artifact your consumer actually needs.
The same documentation warns that the structured evaluator performance-log format can change without notice and recommends codeql generate log-summary rather than direct parsing of that log. “Structured” is not a universal stability guarantee. Consult the documentation for the installed version and avoid introducing an unrelated output migration into a prefix-only repair.
A copyable before-and-after acceptance card
- Wrapper and owner: _____
- Source revision and controlled inputs: _____
- Before / candidate CLI versions: _____ / _____
- Command and output consumer: _____
- Old unlabelled / new labelled fixture: _____ / _____
- Warning / error / unrecognized line handling: _____
- Process exit status captured separately: _____
- Expected artifact exists and consumer accepts it: _____
- Unexpected differences and reviewer: _____
- Decision and rollback reference: _____
For example, a hypothetical wrapper may miss a warning because it expects the message body at column one. The fixture can confirm that matching defect without proving a real scan works. After repairing the wrapper, retain a representative warning case, error case and unrecognized case in its tests. Accept the change only after the actual controlled run and the downstream consumer behave as intended. If a result is missing or an unfamiliar line changes the decision, leave the review unresolved.

For the separate question of whether a scheduled scan actually ran, use the analysis-history triage worksheet. Keep that availability question separate from this output-consumer check. A tidy console deserves a small smile, not an automatic security certificate.