Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 37 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,43 @@ concurrency:
cancel-in-progress: true

jobs:
rust-bridge:
name: Rust bridge (${{ matrix.os }}, ${{ matrix.python-version }})
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
include:
- os: ubuntu-latest
python-version: "3.10"
- os: ubuntu-latest
python-version: "3.14"
- os: macos-latest
python-version: "3.12"
- os: windows-latest
python-version: "3.12"
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: ${{ matrix.python-version }}
- name: Build and install the wheel with native extra
shell: bash
run: |
python -m pip install build
python -m build --wheel
python -c "import glob, subprocess, sys; wheel = glob.glob('dist/*.whl')[0]; subprocess.check_call([sys.executable, '-m', 'pip', 'install', wheel + '[test,cli,rust]'])"
- name: Verify installed package outside checkout imports
run: python -I scripts/check_rust_install.py
- name: Test compatibility, routing, preview and native parity
run: python -m pytest tests/test_contract_481.py tests/test_rust_backend.py tests/test_api_bridge_49.py tests/test_rust_contract.py -q
- name: Generate native parity report
run: python -m tests.rust_contract --output rust-parity.json
- uses: actions/upload-artifact@v4
with:
name: rust-parity-${{ matrix.os }}-${{ matrix.python-version }}
path: rust-parity.json

lint:
runs-on: ubuntu-latest
steps:
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@ docs/*
!docs/make.bat
!docs/optional-surfaces.rst
!docs/migration-4.8.1-contract.md
!docs/migration-4.9.md
!docs/agents/
!docs/agents/**
!docs/audit/
Expand Down
18 changes: 18 additions & 0 deletions CHANGELOG.MD
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,24 @@

## [Unreleased]

#### Added

- Experimental `datafog[rust]` detection with keyword-only `backend="rust"`
on scan/redact entry points; Python remains the default. Core 0.3.1 is pinned,
German requests are explicitly unsupported, and ML composition is unchanged.
- `datafog.compat.v4` preserves the existing scan/redact facade and result classes;
`datafog.v5` previews actual Core types and transformation APIs.
- Native parity tests, installed-wheel CI checks, and public-call benchmarks.

#### Deprecated

- `detect()` and `process()` are now scheduled for removal in 5.0. This explicitly
revises the earlier promise to retain these shims throughout 5.x. They continue
working in 4.9; migrate to scan/redact and review transformation differences.
- OCR/image and Spark/distributed surfaces remain available in 4.9 with use-time
notices and are scheduled for removal in 5.0. Users needing those features can
remain on the final 4.x release.

#### Fixed

- Allow installation on Python 3.14 and certify the core SDK and CLI with
Expand Down
429 changes: 429 additions & 0 deletions GERMAN-PII-CORE-REQUIREMENTS.md

Large diffs are not rendered by default.

145 changes: 145 additions & 0 deletions PLAN-4.9.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
# DataFog 4.9: incremental Core migration

## Outcome

Deliver a bridge release that lets users adopt Rust detection and the Core schema
without changing the default Python behavior. Keep the published 4.8.1 contract
as immutable evidence, with narrowly documented 4.9 API additions and revised
deprecation messages. This work does not publish a release or perform the 5.0
cutover.

## Release decisions

- Keep `pip install datafog` and existing top-level imports working in 4.9.
- Add optional `datafog[rust]`, pinned to the published Core version verified by
this change. Loading the base package must not import the native extension.
- Add keyword-only `backend="python"` to scan/redact entry points. Rust is an
explicit, experimental alternative for `engine="regex"` only.
- Introduce `datafog.compat.v4` for existing scan/redact APIs and result classes;
top-level scan/redact delegate there while preserving class identity.
- Introduce `datafog.v5` as a preview of Core's actual public types and functions,
without translating them into legacy result objects or strategies.
- Revise `detect()` and `process()` warnings: removal is planned for 5.0, replacing
the earlier promise to retain them throughout 5.x. Explain the change publicly.
- Deprecate OCR and Spark in 4.9; remove them in 5.0. Preserve existing optional
installations and behavior during 4.9, and warn at meaningful use sites.
- Leave ML engines, smart/auto composition, CLI text commands, application
integrations, and service APIs on their existing backend in this increment.
- Keep the existing package version until release preparation; this branch
implements 4.9 behavior but does not publish or stamp a final release.

## Increment 1: opt-in Rust detection

1. Extend `datafog.engine.scan` and `scan_and_redact` with keyword-only backend
selection, preserving original parameters, validation and Python defaults.
2. Lazily call the installed `datafog_core.scan`; convert findings to existing
Entity objects using code-point offsets, not byte offsets. Preserve document
order, original text, and legacy regex provenance/confidence conventions.
3. Keep Python allowlist validation/filtering, aliases and transformation logic.
Do not reimplement detectors in the adapter or invoke Core transformations.
4. Reject Rust plus non-regex engines, unsupported German locale/entity requests,
and invalid backend values explicitly. Missing native dependencies must give
an actionable installation error. Never silently retry with Python or return
an empty result on native failure.
5. Explicit-entity redaction must not scan or import the native extension.
6. Add focused tests for routing, Unicode, validation, selection, overlaps,
allowlists, replacement strategies, missing dependencies and propagated errors.

## Increment 2: compatibility namespace and schema preview

1. Make `datafog.compat.v4` the facade for existing scan/redact return shapes and
result classes; retain legacy engine implementation in place where relocation
would break internal imports or monkeypatching contracts.
2. Keep top-level functions and types available and forward backend selection.
Existing agent convenience functions already accept forwarding keyword args.
3. Export the tested Core API through `datafog.v5`, preserving native type
identity. No Core import on plain `import datafog`; clear optional-dependency
errors when the preview is used without the extra.
4. Update detect/process warnings and test the explicit revised removal policy.
Do not expose these functions through the new preview.
5. Test import isolation, facade identity, unchanged legacy results, and real
Core scanning/transformation through the preview.

## Increment 3: retirement notices for optional legacy surfaces

1. Warn on meaningful OCR/Spark use, including direct supported entry points,
without importing heavy dependencies merely to issue a warning.
2. Keep those extras and functionality available throughout 4.9. Document removal
in 5.0 and the option to remain on the final 4.x release. Do not create new
replacement packages as part of this work.
3. Add tests proving notices are emitted and core text imports stay unaffected.

## Increment 4: integration and evidence

1. Keep the 4.8.1 fixture unchanged. Represent approved signature additions and
warning-message changes explicitly in the 4.9 contract checker; compare all
other observed behavior exactly. No blanket skips or recapture from dev.
2. Run applicable cases through the real published Core wheel and produce a
per-case parity report. Separate matches, deliberate unsupported requests,
known detector differences and regressions. Experimental Rust detection must
not be advertised as a drop-in equivalent before gaps are closed.
3. Add CI jobs with and without the Rust extra; smoke-test installed wheels and
native preview behavior on supported Python/platform combinations.
4. Benchmark the public Python and Rust-backed APIs, including conversion and
cold startup, with short, mixed and large synthetic inputs. Keep results
reproducible and make no unsupported speedup claim.
5. Update README, migration documentation and release notes for the actual scope.
6. Run focused suites, broader core regression tests, applicable benchmarks and
pre-commit checks; open a PR against dev and resolve CI failures.

## Work ownership

- Backend agent: engine backend adapter and focused backend tests.
- API agent: compatibility namespace, preview API, top-level delegation and
detect/process retirement warnings, with focused tests.
- Retirement agent: OCR/Spark notices and related tests/documentation.
- Coordinator: packaging, contract integration, parity report, CI, benchmarks,
cross-agent review, final verification and PR.

Agents share one feature branch and have disjoint file ownership. They must not
commit, push, merge, overwrite another agent's files or edit frozen fixtures.
The coordinator reviews and integrates each change before committing.

## Completion criteria

- Default legacy behavior is unchanged except for the documented revised notices.
- Rust use is explicit, installation is optional, failures are visible, and
coverage limitations are documented and exercised by tests.
- Compatibility and native-preview APIs coexist without duplicated native builds.
- OCR/Spark and detect/process have actionable 5.0 retirement notices.
- Tests, installed-wheel checks, parity evidence and benchmarks are reproducible.
- A reviewed, passing PR is ready for dev; merging is subject to user direction
and existing repository approval rules.

## Confirmed scope and delivery

The user confirmed retaining deprecated OCR/Spark functionality in 4.9 and
removing it in 5.0. The coordinator is authorized to merge after checks and
required approvals pass; repository protections remain in effect.

German detector expansion is specified separately in
`GERMAN-PII-CORE-REQUIREMENTS.md`. Implementing or publishing those Core changes
is outside this Python bridge increment. Rust calls requiring that unavailable
coverage must continue to fail explicitly until a tested Core release supports it.

## Implementation evidence

All four increments are implemented on `feature/4.9-core-migration`. Backend,
API, and retirement changes were delegated with disjoint ownership; cross-review
also corrected Windows UTF-8 fixture handling and default-filter CLI notice
visibility. The frozen 4.8.1 fixture is unchanged.

- Base/CLI regression run: 782 passed, with expected optional-dependency skips
and pre-existing corpus xfails.
- Focused native integration run: 347 passed against published Core 0.3.1.
- Python 3.10 and 3.14 base-only runs: 181 passed each; native/CLI-specific tests
skip when their optional dependencies are absent.
- Clean installed-wheel smoke test: compatibility facade, native detection,
native preview, Unicode offsets and transformation all passed.
- Native parity: 61 exact matches, two reviewed detector differences, 17 explicit
unsupported German requests, 31 cases outside this backend's scope.
- Reproducible local timings are in `benchmarks/results-4.9.json`; they include
public-call overhead and cold startup, not just Rust scanning time.

CI and required PR approval remain the final merge gates. No package publication
or Core-repository implementation is part of this delivery.
11 changes: 8 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -188,9 +188,10 @@ scan/redact helpers, or guardrail helpers.
model.
- A Java runtime is required by PySpark.

OCR and Spark are not deprecated. Their broader API and packaging overhaul is
deferred; the 4.x goal is to keep them explicit, documented, and isolated from
the lightweight core path.
The upcoming 4.9 release deprecates OCR and Spark with visible use-time
warnings; their APIs and extras will be removed in 5.0. They remain functional
in 4.9. Users who need these features can stay on the final 4.x release. See
the [4.9 migration guide](docs/migration-4.9.md) for the transition plan.

## Backward-Compatible APIs

Expand Down Expand Up @@ -256,6 +257,10 @@ Telemetry does not include input text or detected PII values.

## Development

The [4.9 migration guide](docs/migration-4.9.md) explains opt-in Rust detection,
the native `datafog.v5` preview, and the revised 5.0 retirement schedule for
`detect`/`process`, OCR, and Spark. The Python detector remains the default.

The [4.8.1 compatibility contract](docs/migration-4.8.1-contract.md) records
published Python behavior for the Rust migration, with frozen fixtures and
instructions for independently reproducing them from the release wheel.
Expand Down
92 changes: 92 additions & 0 deletions benchmarks/compare_detection_backends.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
"""Measure public Python/Rust bridge calls, including adaptation and cold startup."""

import argparse
import importlib.metadata
import json
import os
import platform
import statistics
import subprocess
import sys
import time
from pathlib import Path

os.environ["DATAFOG_NO_TELEMETRY"] = "1"

PAYLOADS = {
"short": ("Contact alice@example.com", 200),
"mixed": (
"Email alice@example.com or call (555) 123-4567. "
"SSN 123-45-6789; card 4111 1111 1111 1111; host 192.168.1.1.",
100,
),
"large_sparse": ("ordinary prose " * 70_000 + " alice@example.com", 3),
}


def measure(fn, text, backend, loops, samples):
result = fn(text, backend=backend)
durations = []
for _ in range(samples):
start = time.perf_counter()
for _ in range(loops):
fn(text, backend=backend)
durations.append((time.perf_counter() - start) / loops * 1_000_000)
return {
"median_us": statistics.median(durations),
"samples_us": durations,
"entities": len(result.entities),
"iterations_per_sample": loops,
}


def cold_start(backend, samples):
source = (
"import datafog; "
f"datafog.scan('Contact alice@example.com', backend={backend!r})"
)
durations = []
for _ in range(samples):
start = time.perf_counter()
subprocess.run([sys.executable, "-c", source], check=True, capture_output=True)
durations.append((time.perf_counter() - start) * 1_000)
return {"median_ms": statistics.median(durations), "samples_ms": durations}


def main():
import datafog

parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--samples", type=int, default=5)
parser.add_argument("--output", type=Path, required=True)
args = parser.parse_args()
if args.samples < 1:
parser.error("samples must be positive")
report = {
"python": platform.python_version(),
"platform": platform.platform(),
"core_version": importlib.metadata.version("datafog-core"),
"method": "One warmup call, median repeated calls; no ML. Cold = process+import+first scan.",
"warm": [],
"cold": {},
}
for name, (text, loops) in PAYLOADS.items():
for operation in ("scan", "redact"):
row = {
"payload": name,
"operation": operation,
"utf8_bytes": len(text.encode()),
}
for backend in ("python", "rust"):
row[backend] = measure(
getattr(datafog, operation), text, backend, loops, args.samples
)
report["warm"].append(row)
for backend in ("python", "rust"):
report["cold"][backend] = cold_start(backend, args.samples)
args.output.write_text(json.dumps(report, indent=2) + "\n")
print(f"Saved end-to-end measurements to {args.output}")


if __name__ == "__main__":
main()
Loading
Loading