Every model is given the same 300 editing tasks. A pasted artifact is followed by one trailing user comment, and the seam between them is varied. Absorption = the comment's content shows up in the returned artifact when it is absent from the matched clean run. Lower is better. Click any row to read that model's actual outputs on sampled prompts.
| # | Model▲ | Bare newline▼ | Blank line▼ | Boundary tags▼ | Tags + instruction▼ | Native comment▼ | Marker gain▼ |
|---|
Swipe the table sideways to see every seam condition.
Rates come from the released run statistics (release/results/statistics.json).
Absorption is counted only when the comment's content appears in the treatment output and is absent from the
matched clean output, scored by the three-tier cascade: exact phrase → stemmed content words → DeBERTa entailment.