fix(responses): skip structured parsing for commentary (#3861)
- [x] I understand that this repository is auto-generated and my pull request may not be merged ## Changes being requested Parse structured output only for messages with `phase="final_answer"` or a missing/null phase. Commentary and other explicit phases retain their text and metadata with `parsed=None`. Currently, a response containing prose commentary followed by a valid final JSON message raises a Pydantic validation error while parsing the commentary. If the commentary happens to be schema-valid JSON instead, `output_parsed` returns that intermediate value rather than the final answer. The phase policy is centralized in `parse_text()` in handwritten `lib/` code. Both `parse_response()` and the streaming `response.output_text.done` handler pass the message phase to it; stream completion already uses `parse_response()`. `helpers.md` documents the behavior, including forward compatibility for unknown explicit phases. With commentary unparsed, the existing `output_parsed` accessor reaches the final result without modification. Messages with `phase="final_answer"`, `phase=None`, or an omitted phase retain existing behavior, including validation errors for invalid final JSON. Commentary text remains available through normal text fields and events. The request's strict JSON Schema, refusals, and tool-call handling are unchanged. ## Additional context & links Fixes #3859. Requesting SDK CODEOWNER review for this structured-parsing behavior change. ### Validation At current head `9594078438efabc943a1d0da79e46f209ea3c106`, after the maintainer update: - Responses tests: **229 passed with Pydantic v1** (Python 3.13.5) and **229 passed with Pydantic v2** (Python 3.10.20). - Added coverage includes unknown explicit phases, legacy streaming phases, multiple text parts, and intermediate messages before/after final output. The following baseline reproduction and full-suite results were obtained at the original PR head `f3cb72e3de1e4f1753002e362dcb2d45603d2762`: Based on `main` at `d421d7ab8c0a5e4e00147407ef9941c7f2bc9c09`. The regression tests use synthetic JSON responses and SSE through `httpx2.MockTransport` and the public synchronous/asynchronous `responses.parse()` and `responses.stream()` entry points. They require no API credentials or live model calls. Every structured request also checks that the strict JSON Schema is still sent. Coverage includes prose and schema-valid commentary, absent/null phases, commentary without a final result, final refusals, invalid final JSON/schema, streamed text-done and completion parsing, interleaved tool calls, and null/missing/empty completion output. - New regression tests: **34 failed / 24 passed on unchanged main → 58 passed with this fix**, using identical fixtures. - Full SDK suite on Python 3.13.5 (`./scripts/test -n 4`): **13,018 passed / 162 skipped with Pydantic v2**, and **13,004 passed / 176 skipped with Pydantic v1**. - Responses tests: **165 passed with Pydantic v1**; **165 passed on Python 3.10.20 with Pydantic v2**. - `./scripts/lint` on Python 3.10.20: Ruff, Pyright, Mypy, and import check passed. - Ruff formatting for changed files and `git diff --check` passed. - The custom-code budget check using main's unchanged checker passed; no budget or generation metadata changes. --------- Co-authored-by: naotaka1128 <5688448+naotaka1128@users.noreply.github.com> Co-authored-by: Marcus Wood <marcuswood@openai.com>
M
ML_Bear committed
6520df1220a0030c7b45b00a6fa4bf2e9c7396ac
Parent: b3c04f4
Committed by GitHub <noreply@github.com>
on 9/15/2026, 9:55:19 PM