SIGN IN SIGN UP

fix: stop copying a chunk's orig_elements twice on every serialization (#4472)

## What & why

**Problem:** Anyone chunking a document was paying to photocopy the
whole thing twice on every serialization, then bin both copies.
`ElementMetadata.to_dict()` deep-copies every metadata field and then
replaces `coordinates`, `data_source`, `orig_elements` and
`key_value_pairs` with their serialized form, so the copies of those
four are built and thrown away. Separately,
`_fix_metadata_field_precision()` copies every element in order to round
`coordinates` and `detection_class_prob`, which most elements do not
have. On a chunk, `orig_elements` holds every source element of that
chunk, so a single `to_dict()` duplicated the document twice over. In a
profiled local pipeline over 45,000 elements, `copy.deepcopy` and its
helpers accounted for roughly 40% of total run time.

There is a correctness consequence too. `Element.id` mints a uuid on
first access and caches it on that element. Because the copies were the
objects that got serialized, they took the freshly minted ids with them
and the originals stayed unset, so serializing one chunk twice reported
different `element_id` values for the same source elements each time.

**Change:** In `to_dict()`, deep-copy only the fields that are not
separately serialized and rebuild the dict in its original key order, so
the four re-serialized fields keep their position. In
`_fix_metadata_field_precision()`, return the element untouched when it
has neither `coordinates` nor `detection_class_prob`; mint the element's
id before the copy that remains; and make that copy shallow, rebuilding
`CoordinatesMetadata` instead of rounding the caller's points in place,
so `orig_elements` is not duplicated and its nested ids stay stable.

**Blast radius:** 3/5 -- two functions on the shared serialization path
that every caller of `to_dict()`, `elements_to_json()`,
`elements_to_dicts()` and `orig_elements` inherits; small,
self-contained, and revert-safe.

## Linked ticket

none

## Impact

**Library users:** chunk-heavy serialization gets materially faster, and
repeated serialization of the same element now reports stable
`element_id` values for its `orig_elements`. Measured on this branch,
one probe run in both states, interleaved in a single process, minimum
of N:

| case | before | after | speedup |
|---|---:|---:|---:|
| `to_dict`, no `orig_elements` | 17.2 ms | 18.3 ms | 0.94x |
| `to_dict`, 10 `orig_elements` | 323.1 ms | 74.2 ms | 4.36x |
| `to_dict`, 40 `orig_elements` | 399.1 ms | 62.4 ms | 6.40x |
| `elements_to_json`, no `orig_elements` | 77.9 ms | 22.2 ms | 3.51x |
| `elements_to_json`, 10 `orig_elements` | 440.6 ms | 75.3 ms | 5.85x |
| `elements_to_json`, 40 `orig_elements` | 666.1 ms | 63.0 ms | 10.57x |

The first row is the one that did not improve and reads slightly worse:
it pays four extra dict pops and has no `orig_elements` to skip. The
machine was under heavy load during timing, so treat the magnitudes as
approximate and that row as indistinguishable from noise.

**Wire contract / clients:** the serialized dict is unchanged in
structure and in every value except `element_id` for elements that had
none assigned, which was previously regenerated on each call. This
reaches `elements_to_json()` and `elements_to_ndjson()` as well as the
ids inside `orig_elements`. A caller that recorded those ids and
expected a later serialization to produce the same ones was already
getting different values every time; it now gets the same ones. This
holds for an element that carries `orig_elements` together with its own
`coordinates` or `detection_class_prob` as well, because the precision
fix no longer deep-copies the `orig_elements` subtree. Elements given an
explicit `element_id`, or one assigned by `id_to_hash()`, were never
affected either way.

Metadata key order is preserved. An earlier revision of this branch
popped the four separately-serialized fields before the copy and
appended them afterwards, which moved `coordinates`, `data_source`,
`orig_elements` and `key_value_pairs` to the end of the dict.
`elements_to_json()`, `elements_to_ndjson()` and the base64
`orig_elements` all serialize with `sort_keys=True` so their bytes never
changed, but `Element.to_dict()` and `elements_to_dicts()` are
order-sensitive for a caller that writes them unsorted, and
unstructured-ingest's chunker does exactly that. That turned
`test_ingest_src` red on a fixture diff. The current revision rebuilds
the dict in the original order and a test pins it.

Shared callers that inherit this: `ElementMetadata.to_dict()` is reached
from `Element.to_dict()`, and therefore from `elements_to_dicts()` (and
its `convert_to_isd` / `convert_to_dict` aliases), `elements_to_json()`,
`elements_to_base64_gzipped_json()` and `elements_to_ndjson()`.
`_fix_metadata_field_precision()` is called by
`elements_to_base64_gzipped_json()` (`base.py:256`),
`elements_to_json()` (`base.py:453`) and `elements_to_ndjson()`
(`base.py:475`).

## Risk / rollback

Low. Two functions, no signature or schema change, revert the commit to
back it out. The one deliberate behavior change is the `element_id`
stability described above.

## How it was verified

Ran locally on Python 3.13 against this branch. Each of the three new
tests was run first against the unfixed sources restored from `HEAD`
with the tests in place, to confirm it fails for the reason claimed, and
then against the fix.

Suites run: `test_unstructured/chunking`, `test_unstructured/documents`
and `test_unstructured/staging` give 796 passed and 25 skipped on the
current head, with one collection error in
`test_unstructured/staging/test_huggingface.py` for a missing
`transformers`, which is an uninstalled optional dependency rather than
a failure. `make check-ruff` is clean. The wider `test_unstructured`
tree passes apart from the `partition` and `metrics` trees, which need
`unstructured_inference`, and `cleaners/test_translate.py` plus the
benchmark test, which fail on missing `sentencepiece` and
`pytest-benchmark` in my environment and fail the same way without this
change.

Note on the environment, because it bit me: without the `csv`, `docx`,
`tsv` and `xlsx` extras installed,
`test_unstructured/staging/test_base.py` and
`test_unstructured/chunking/test_basic.py` do not collect at all, and
the run reports a confident 564 passed while silently skipping the file
the precision tests live in.

Reviewed by fable and GPT-5.5 Pro before this leaves draft. Both
independently found that the first commit left ids unstable for any
element carrying `coordinates` or `detection_class_prob`, which is every
hi_res-partitioned element, so its stability claim was false for the
common case. My own test could not see it, because I built the fixture
from bare `Text` elements with no coordinates. Fixed, with both variants
now covered by a parametrized regression test. fable separately caught
that the `CHANGELOG.md` heading has to carry the `-dev0` suffix to match
`__version__` or `scripts/version-sync.sh` fails `make check`: the
precedent is commit `4fe4097`, whose heading is `## 0.27.5-dev0`, and
the release commit is what strips the suffix from both.

Not verified: `scripts/version-sync.sh -c` could not run locally, since
it needs GNU sed 4.3 and macOS ships BSD sed, so CI is the first real
check that the heading and `__version__` agree. I also have not measured
this on a GPU or OCR-heavy end-to-end partition, where model inference
dominates and this saving is proportionally much smaller.

## Proof

**Repro.** Profiled a 201-document, 45,000-element local pipeline with
the real chunker under cProfile. `copy.deepcopy` was 10.783 s exclusive
over 180,000 calls in a 40.018 s run, and with `_deepcopy_list`,
`_deepcopy_dict`, `_keep_alive` and `_deepcopy_atomic` the deepcopy
machinery totalled about 16.6 s. Reduced to a standalone reproduction of
the id half:

```
$ python -c "from unstructured.documents.elements import Text, ElementMetadata; \
    md = ElementMetadata(orig_elements=[Text('hello world')]); e = Text('chunk', metadata=md); \
    print(e.to_dict() == e.to_dict())"
False
```

Decoding the base64 payload on three consecutive calls gave three
different `element_id` values for the same nested element:
`2ba29618-...`, `4e806db8-...`, `7070f009-...`.

**Failing tests, run against the unfixed sources with the new tests in
place.**

```
$ python -m pytest test_unstructured/documents/test_elements.py -k "does_not_deep_copy or same_way_on_every_call" test_unstructured/staging/test_base.py -k "does_not_deep_copy or same_way_on_every_call or leaves_elements_with_nothing_to_round" -q
FAILED test_unstructured/documents/test_elements.py::DescribeElementMetadata::and_it_does_not_deep_copy_the_sub_objects_it_reserializes
FAILED test_unstructured/documents/test_elements.py::DescribeElementMetadata::and_it_serializes_orig_elements_the_same_way_on_every_call
FAILED test_unstructured/staging/test_base.py::test_fix_metadata_field_precision_leaves_elements_with_nothing_to_round_alone
3 failed, 145 deselected in 1.71s
```

**After the fix, same tests, same command:**

```
2 passed, 64 deselected in 2.41s     (documents)
1 passed, 81 deselected in 1.31s     (staging)
```

**Second round, after review.** With the id-mint line removed from
`staging/base.py`:

```
$ python -m pytest test_unstructured/documents/test_elements.py -k "precision_still_has_to_be_rounded" -q
FAILED ...and_that_holds_for_elements_whose_precision_still_has_to_be_rounded[coordinates]
FAILED ...and_that_holds_for_elements_whose_precision_still_has_to_be_rounded[detection_class_prob]
2 failed, 66 deselected in 2.24s
```

With it restored: `2 passed, 66 deselected in 1.33s`. All three element
kinds checked directly:

```
with coordinates, two to_dict calls equal: True
with detection_class_prob, two to_dict calls equal: True
no coordinates, two to_dict calls equal: True
```

**Suites for the touched areas:**

```
$ python -m pytest test_unstructured/chunking test_unstructured/documents test_unstructured/staging -q
673 passed, 25 skipped in 13.18s
```

**The differential probe, one command run in both states.** This is the
table under Impact; the row that did not move is the point of showing
all six.

```
case                       baseline        fixed   speedup
to_dict, no origs           17.2 ms      18.3 ms     0.94x
to_dict, 10 origs          323.1 ms      74.2 ms     4.36x
to_dict, 40 origs          399.1 ms      62.4 ms     6.40x
to_json, no origs           77.9 ms      22.2 ms     3.51x
to_json, 10 origs          440.6 ms      75.3 ms     5.85x
to_json, 40 origs          666.1 ms      63.0 ms    10.57x
```

**Still shaky.** The timings were taken under load average 55, so their
direction and rough scale are solid and the precise multiples are not.
The 0.94x row cannot be separated from noise at that load.

**Third round, after review.** Three more tests, each run against the
unfixed sources with the test in place.

Key order, `test_unstructured/documents/test_elements.py`:

```
FAILED ...and_it_emits_the_separately_serialized_fields_in_their_declared_position
E       At index 0 diff: 'filetype' != 'coordinates'
```

Nested id stability under the precision fix,
`test_unstructured/staging/test_base.py`, both variants failing at the
old head with differing uuids across two `elements_to_json()` calls:

```
FAILED ...test_elements_to_json_keeps_orig_elements_ids_stable_for_an_element_it_rounds[coordinates]
FAILED ...test_elements_to_json_keeps_orig_elements_ids_stable_for_an_element_it_rounds[detection_class_prob]
```

All three pass after the fix.

The third new test,
`test_fix_metadata_field_precision_rounds_a_copy_and_leaves_the_callers_element_unrounded`,
passes at both heads, because the old `deepcopy` already protected the
caller. A test that cannot fail is worth nothing, so it was
mutation-checked instead: with the shallow copy plus in-place rounding
of `points` it fails on `(1.23, 2.35) != (1.23456, 2.34567)`.

**Not re-measured.** The timing table above was taken before the
key-order rebuild was added. That rebuild is one extra dict
comprehension over the same fields, measured at 1.47 ms against 1.42 ms
on a 40-element chunk, so the table's magnitudes still stand, but it is
not a fresh run.

**Not run.** The ingest shell fixtures (`test-ingest-src.sh`) were not
run locally. The key-order claim rests on a direct `to_dict()` probe
against `origin/main` across three metadata shapes, not on the CI
fixture diff, so CI is the first real check of it.

## Dependencies / merge order

none


<!-- This is an auto-generated description by cubic. -->
<a
href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4472?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>
<!-- End of auto-generated description by cubic. -->

---------

Co-authored-by: paulkarayan <pk@unstructured.io>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
P
Paul Karayan committed
2c0c7a6ca71334bdfbbcb083311c4c859c0b59c7
Parent: 0242693
Committed by GitHub <noreply@github.com> on 10/3/2026, 4:03:08 PM