SIGN IN SIGN UP

fix(responses): skip structured parsing for commentary (#3861)

- [x] I understand that this repository is auto-generated and my pull
request may not be merged

## Changes being requested

Parse structured output only for messages with `phase="final_answer"` or
a missing/null phase. Commentary and other explicit phases retain their
text and metadata with `parsed=None`.

Currently, a response containing prose commentary followed by a valid
final JSON message raises a Pydantic validation error while parsing the
commentary. If the commentary happens to be schema-valid JSON instead,
`output_parsed` returns that intermediate value rather than the final
answer.

The phase policy is centralized in `parse_text()` in handwritten `lib/`
code. Both `parse_response()` and the streaming
`response.output_text.done` handler pass the message phase to it; stream
completion already uses `parse_response()`. `helpers.md` documents the
behavior, including forward compatibility for unknown explicit phases.

With commentary unparsed, the existing `output_parsed` accessor reaches
the final result without modification. Messages with
`phase="final_answer"`, `phase=None`, or an omitted phase retain
existing behavior, including validation errors for invalid final JSON.
Commentary text remains available through normal text fields and events.
The request's strict JSON Schema, refusals, and tool-call handling are
unchanged.

## Additional context & links

Fixes #3859.

Requesting SDK CODEOWNER review for this structured-parsing behavior
change.

### Validation

At current head `9594078438efabc943a1d0da79e46f209ea3c106`, after the
maintainer update:

- Responses tests: **229 passed with Pydantic v1** (Python 3.13.5) and
**229 passed with Pydantic v2** (Python 3.10.20).
- Added coverage includes unknown explicit phases, legacy streaming
phases, multiple text parts, and intermediate messages before/after
final output.

The following baseline reproduction and full-suite results were obtained
at the original PR head `f3cb72e3de1e4f1753002e362dcb2d45603d2762`:

Based on `main` at `d421d7ab8c0a5e4e00147407ef9941c7f2bc9c09`.

The regression tests use synthetic JSON responses and SSE through
`httpx2.MockTransport` and the public synchronous/asynchronous
`responses.parse()` and `responses.stream()` entry points. They require
no API credentials or live model calls. Every structured request also
checks that the strict JSON Schema is still sent.

Coverage includes prose and schema-valid commentary, absent/null phases,
commentary without a final result, final refusals, invalid final
JSON/schema, streamed text-done and completion parsing, interleaved tool
calls, and null/missing/empty completion output.

- New regression tests: **34 failed / 24 passed on unchanged main → 58
passed with this fix**, using identical fixtures.
- Full SDK suite on Python 3.13.5 (`./scripts/test -n 4`): **13,018
passed / 162 skipped with Pydantic v2**, and **13,004 passed / 176
skipped with Pydantic v1**.
- Responses tests: **165 passed with Pydantic v1**; **165 passed on
Python 3.10.20 with Pydantic v2**.
- `./scripts/lint` on Python 3.10.20: Ruff, Pyright, Mypy, and import
check passed.
- Ruff formatting for changed files and `git diff --check` passed.
- The custom-code budget check using main's unchanged checker passed; no
budget or generation metadata changes.

---------

Co-authored-by: naotaka1128 <5688448+naotaka1128@users.noreply.github.com>
Co-authored-by: Marcus Wood <marcuswood@openai.com>
M
ML_Bear committed
6520df1220a0030c7b45b00a6fa4bf2e9c7396ac
Parent: b3c04f4
Committed by GitHub <noreply@github.com> on 9/15/2026, 9:55:19 PM