Establish the published 4.8.1 Python text API contract - #177
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Rust migration needs a reproducible baseline of what the released Python package actually does. This freezes 111 black-box observations from the published
datafog==4.8.1wheel and checks the development checkout against them without regenerating expectations.The cases cover structured synchronous detection, Unicode offsets, selection and aliases, allowlists, German locale activation, explicit-span overlap handling, all legacy transformation strategies, result shapes, selected signatures, warnings/errors, convenience APIs, guardrails, and regex TextService calls. The migration guide separates 102 compatibility cases from nine observations requiring review, including invalid-checksum cards and mismatched explicit span text. ML, OCR, Spark, CLI, async, and application-adapter acceptance remain explicitly outside this baseline's scope.
The capture tool requires isolated imports, verifies the pinned PyPI wheel SHA-256 and installed package bytes, and writes only a separate candidate file. The existing CI test discovery automatically includes the contract suite. Detection and transformation behavior are unchanged. The
nlp-advancedextra now explicitly includes SentencePiece and protobuf so the multilingual GLiNER tokenizer can load without the OCR extra. Its install-profile smoke test also checks that the tokenizer protobuf schema loads.Validation: