Before upgrading Python, identify who produces each text file or stream and who reads it next. A default encoding can change while the bytes at either end stay exactly the same.
Python 3.15.0rc3 arrived on October 2, 2026. As checked October 7, it remains a preview that the release team does not recommend for production. The release schedule lists October 9 as the expected final-release date. This is a useful testing window for one specific change: Python 3.15 enables UTF-8 mode by default.

Separate the default from the data contract
PEP 686 changes default text encoding behavior for files, standard streams and pipes. It also retains the option to disable UTF-8 mode. Existing programs that depend on a different default can encounter decoding errors or corrupted-looking text.
The practical implication is that upgrading an interpreter does not transcode an old export or tell another program to emit different bytes. Review the agreement at each handoff. An explicit codec should match the producer's documented format; an unknown format needs investigation rather than a guess.
For a small illustration, the character é is hexadecimal C3 A9 in UTF-8 and E9 in Windows-1252, also called CP1252. These byte values can represent the same intended character under different codecs. They are not interchangeable inputs to a UTF-8 decoder.

Boundary one: a file enters Python
Start with one important import: a supplier export, a saved report, or an application configuration file. Record the producing application and its export setting. Keep a synthetic sample with known expected text, including characters beyond plain ASCII.
Python's text I/O reference documents the default and the optional EncodingWarning. Review open(), Path.read_text() and wrappers around them. When the file contract specifies UTF-8, make that intent explicit with encoding="utf-8". A known CP1252 file needs its own matching codec.
Do not infer a format from a filename extension or one successful decode. Ask for the producer's setting or specification and compare the decoded value with a known expectation. Keep original bytes intact while investigating; replacing unreadable characters can hide the evidence you need.
Boundary two: a command sends bytes to Python
The subprocess reference distinguishes binary streams from text mode. With text=True, or a supplied encoding or errors argument, Python uses text wrappers. Without text mode, captured output remains bytes.
For each command you capture, write down how that command chooses its output encoding. Setting encoding="utf-8" on the Python call chooses Python's decoder; it does not configure the external producer. If the command has a documented output-format setting, record it separately.
Test a representative non-ASCII result and any error output your application consumes. If the producer contract is unresolved, retain a sanitized byte fixture and assign an owner to establish the format. Do not treat a decode-error suppression flag as a successful migration.
Boundary three: Python writes an export
Identify the receiving application before changing a writer. A generated file may be valid UTF-8 while its intended reader expects something else. Record the output codec, any documented byte-order-mark requirement, and the receiving application's import setting as separate facts.
Use a disposable output path and reopen the result in the intended consumer. Compare the exact text, not merely whether the file opens. Where byte-level equality matters, check that separately from text equality, since newline conventions can also affect the bytes. Use made-up records instead of customer names or private exports.
Record the environment before comparing modes
The UTF-8 mode reference explains that mode is set at interpreter startup and is visible through sys.flags.utf8_mode. Standard-stream encodings can also be overridden, including through PYTHONIOENCODING. Record the actual interpreter version, operating system, locale and relevant stream encodings rather than assuming every environment behaves alike.

For an existing Python 3.11-or-newer test environment, locale.getencoding() reports the locale encoding independently of UTF-8 mode. An intentional encoding="locale" choice keeps locale dependence; it does not promise the same format on every computer.
- Find exercised defaults. Run representative tests with
-X warn_default_encoding. A warning is a review lead. No warning does not establish coverage of code the tests never reached. - Repeat identical fixtures. Use separate test processes with
-X utf8=0and-X utf8=1, keeping other inputs constant. The command-line reference documents these switches. - Compare expected text. Record exceptions, incorrect characters and downstream readback. A process exit code alone does not answer whether the text survived.
- Repeat on the target runtime. Testing an older interpreter's UTF-8 mode is a focused rehearsal, not complete Python 3.15 compatibility evidence.

A small synthetic check, with a narrow conclusion
A check prepared for this guide ran on Linux with Python 3.12.14, an isolated C locale reporting ANSI_X3.4-1968, and UTF-8 mode both off and on. Its UTF-8 fixture contained café 서울; its CP1252 fixture contained café £. A child process emitted predetermined CP1252 bytes for the pipe case.
Explicit UTF-8 file reading and writing preserved the expected text in both modes. Explicit CP1252 file and pipe decoding also matched in both. Implicit UTF-8 file reading and writing failed with mode off and matched with mode on; implicit CP1252 file and pipe decoding failed in both modes.
This demonstrates why producer-specific contracts matter. It does not simulate a Windows code page, test Python 3.15, or validate a third-party application. The byte comparison in the illustration was also checked directly: "é".encode("utf-8").hex() gives c3a9, while CP1252 gives e9.
Copy the text-boundary worksheet
Use one row per handoff, rather than one overall pass mark for an application. This is an original review template, not a Python release requirement.
| Field | What to record |
|---|---|
| Boundary | File import, command output or export; producer → consumer; responsible owner. |
| Encoding evidence | Producer specification or export setting, checked date, and the relevant codec. Write unresolved if unknown. |
| Current Python call | API or wrapper, current encoding argument, and any error-handling behavior. |
| Intended contract | Explicit named codec, deliberate locale dependence, or a question awaiting its owner. |
| Fixture and expectation | Sanitized input bytes, exact expected text, and expected downstream interpretation. |
| Mode comparison | Interpreter and environment; mode-off and mode-on outcomes; actual text or exception. |
| Consumer check | Receiving application and import setting; observed readback; byte check when required. |
| Decision | Contract verified, change needed, or unresolved; next owner and a repeatable test reference. |

Finish with a short list of resolved contracts and named open questions. If UTF-8 mode must temporarily remain disabled, record why, who owns the dependency, and what test will allow that exception to end. Keep the broader runtime upgrade under the project's normal compatibility and release review.