Skip to content

Rust Engine PyO3 Boundary

This document records the native boundary that backs the current Python Bank wrapper and the CLI/MCP extraction surfaces, including the production matcher and the two internal measurement modes.

Native Bank Constructors

from nerb import _engine

bank = _engine.Bank.from_source_bytes(b'{"CODE":{"Alpha":"A"}}', format_hint="json")
canonical = bank.to_canonical_json_bytes()
round_tripped = _engine.Bank.from_canonical_json_bytes(canonical)

assert round_tripped.metadata()["bank_hash"] == bank.metadata()["bank_hash"]

metadata() exposes only engine and canonical-bank facts needed by Python wrappers:

{
  "engine": "nerb_engine",
  "build_source_sha256": "sha256:...",
  "schema": 1,
  "bank_hash": "sha256:...",
  "entity_count": 1,
  "pattern_count": 1,
  "defaults": {
    "engine": "rust-regex-meta",
    "unicode": true,
    "case_insensitive": false,
    "word_boundaries": false,
    "normalization": "none"
  },
  "compile_options": {
    "match_mode": "entity_independent"
  },
  "match_mode": {
    "name": "entity_independent",
    "status": "production_default",
    "production_default": true,
    "internal_only": false,
    "semantic_notes": "reports cross-entity overlap with leftmost-first matching within each entity"
  },
  "scan_limits": {
    "maximum_input_bytes": 10485760,
    "maximum_concurrent_scans_per_bank": 8
  },
  "regex_resources": {
    "scope": "entity_independent_shards",
    "physical_regex_layers": 0,
    "maximum_regex_layers_per_entity": 0,
    "compiled_regex_static_bytes": 0,
    "eager_cache_bytes_per_scan": 0,
    "pikevm_cache_projection_bytes_per_scan": 0,
    "pikevm_stack_growth_allowance_bytes_per_scan": 0,
    "lazy_dfa_growth_allowance_bytes_per_scan": 0,
    "regex_cache_allowance_bytes": 0,
    "size_limit_bisections": 0,
    "resource_limit_bisections": 0,
    "accounted_bytes": 0,
    "cache_concurrency_budget": 8,
    "explicit_regex_cache_slots": 8,
    "internal_meta_cache_pool_used": false,
    "per_lazy_dfa_cache_capacity_bytes": 32768,
    "maximum_lazy_dfa_caches_per_regex": 3,
    "pikevm_stack_nfa_memory_multiplier": 16,
    "onepass_enabled": false,
    "bounded_backtracker_enabled": false,
    "maximum_layers_per_entity": 128,
    "maximum_patterns_per_regex_layer": 128,
    "maximum_accounted_bytes": 805306368
  },
  "detectors": [
    {
      "detector_index": 0,
      "entity": "CODE",
      "canonical_name": "Alpha",
      "surface_name": "Alpha",
      "stable_id": "pattern:sha256:...",
      "priority": 0
    }
  ]
}

The module-level _engine.BUILD_SOURCE_SHA256 equals each bank's build_source_sha256. The build hashes a closed, sorted inventory containing Cargo.toml, Cargo.lock, build.rs, and every Rust source file after LF normalization; the build fails if that inventory changes without an explicit update. Only production entity_independent metadata contains regex_resources, because its accounting does not describe the internal modes' additional matchers.

MatchBuffer

MatchBuffer is a Rust-owned container for raw scan results:

pub struct RawMatch {
    pub detector_index: u32,
    pub start_byte: u64,
    pub end_byte: u64,
}

Python can create, reserve, clear, and inspect a buffer without constructing public record dictionaries:

buffer = _engine.MatchBuffer(capacity=1024)
assert len(buffer) == 0

raw = _engine.MatchBuffer.from_raw_matches([(7, 10, 15)])
assert raw[0] == (7, 10, 15)
raw.clear()

from_raw_matches accepts a sized Python sequence and exists to test the boundary. Bank.scan_bytes fills MatchBuffer from Rust. Public record projection remains outside the scan loop.

Python-created buffers and Rust scanner appends are capped at 1,000,000 requested raw matches and use fallible Rust allocation paths. Later dense-hit measurement may revisit this logical limit.

Scanning

Bank.scan_bytes implements the production-default entity_independent mode: one logical matcher per entity with leftmost-first semantics inside each entity and cross-entity overlap preserved. A logical matcher may contain bounded exact-literal Aho-Corasick layers, mapped Aho-Corasick layers for supported normalized-whitespace and simple-fold literals, and one or more residual regex-automata layers. Global arbitration by start offset and original pattern order reconstructs the entity's leftmost-first result. It validates UTF-8 input, releases the GIL during the Rust scan, returns raw (detector_index, start_byte, end_byte) matches sorted by byte offsets and detector index, and optionally fills a caller-provided MatchBuffer.

Matcher construction disables one-pass and bounded-backtracker strategies and applies bounded regex-automata NFA, hybrid-cache, and DFA limits. Size-limit failures for advancing residual patterns are deterministically bisected, with explicit per-entity layer and aggregate static/cache-memory ceilings. Unsupported syntax, singleton size failures, or aggregate resource exhaustion fail during bank construction before the bank is returned; syntax/shape failures raise ValueError, while aggregate memory-budget exhaustion raises MemoryError. Production metadata reports the realized layer, memory, and bisection profile. Cache accounting includes every physical regex, including the first or only regex in an entity. NERB does not use the meta regex's internal sharded cache pool: each physical regex owns exactly eight explicit cache slots selected by the bank's scan permit. For one scan slot, the allowance sums the eager meta-cache heap, a checked projection of the equivalent PikeVM fallback's fixed cache, a conservative PikeVM epsilon-stack growth bound, and three 32 KiB lazy-DFA capacities (forward, reverse, and reverse-inner). regex_cache_allowance_bytes multiplies that per-scan sum by eight, and accounted_bytes adds compiled static bytes. One-pass and bounded-backtracker strategies are disabled so no unmeasured lazy strategy cache can appear.

An initial physical regex layer contains at most 128 patterns. This named envelope bounds the PikeVM's implicit-capture state/slot product before compilation while retaining deterministic pattern order. A layer that still exceeds a compile-size or accounted-resource limit is bisected deterministically; metadata reports size-limit and resource-limit bisections separately. More than 128 physical layers in one entity or more than 768 MiB of aggregate static-plus-eight-slot cache allowance fails before the layer is committed.

The PikeVM stack bound is 16 times the compiled implicit-capture NFA's reported memory. In pinned regex-automata 0.4.14, every epsilon-stack push is backed by an NFA state or stored Union alternate; the largest private frame is under four machine words. The multiplier covers the worst frame-to-StateID ratio and geometric Vec retained capacity. The fixed-cache projection uses checked arithmetic for the four NFA-state ID vectors and two capture-slot tables that regex-automata creates. Production accounting therefore rejects arithmetic overflow without first allocating a potentially quadratic table. Safe-size regressions compare the projection with an actual PikeVM cache, including a complex cache larger than 32 KiB, and exercise overflow failure directly.

Every compiled bank admits at most eight scans at once. A per-bank permit covers every native scan mode and is released on success, validation/allocation failure, or panic unwinding; additional callers wait without reducing the first eight to a serial lane. This enforced scan ceiling is the same concurrency value used by regex-cache accounting.

All inline scan variants reject inputs larger than 10 MiB before releasing the GIL or allocating mapped-haystack projections. Exactly 10 MiB is accepted. scan_text inherits this byte limit after UTF-8 encoding, so a Unicode string's encoded length—not its Python character count—is authoritative. Reused MatchBuffer objects are cleared on an over-limit error, just as they are for other scan failures.

IGNORECASE, MULTILINE, DOTALL, and VERBOSE are applied through per-pattern syntax configuration. The ASCII flag lowers ASCII-sensitive escapes and boundaries such as \w, \d, \s, and \b while leaving the rest of the detector pattern in UTF-8-safe Unicode regex mode.

bank = _engine.Bank.from_source_bytes(b'{"PERSON":{"Sam":"Sam"},"PROJECT":{"Samba":"Samba"}}')
raw = bank.scan_bytes(b"Samba ships")
assert [raw[i] for i in range(len(raw))] == [(0, 0, 3), (1, 0, 5)]

All Overlaps Prototype

compile_options_json='{"match_mode":"all_overlaps"}' builds an internal prototype around lower-level regex-automata hybrid DFAs:

  • a forward DFA runs overlapping search with MatchKind::All;
  • a reverse DFA with per-pattern start states recovers the start byte for each reported end;
  • local pattern IDs are translated back to global detector indexes before appending to MatchBuffer.

The prototype rejects Unicode word-boundary assertions such as \b because the lower-level DFA only provides heuristic Unicode-boundary support that can quit on valid non-ASCII UTF-8. Use explicit ASCII word-boundary syntax such as (?-u:\b) for raw all_overlaps, or use the production-default entity_independent mode for Unicode boundary semantics.

Raw all_overlaps output is intentionally not the default contract. It preserves cross-entity overlap, but it also reports within-entity overlapping detectors and every matching span for each detector pattern. It does not preserve a separate branch identity inside one regex; attribution still stops at the NERB detector index. For example, the production entity_independent mode chooses Samwise for Samwise|Sam over Samwise, while raw all_overlaps exposes both (0, 0, 3) and (0, 0, 7) for that one detector. That means a span-only candidate post-filter cannot prove exact leftmost-first reconstruction.

The prototype therefore exposes Bank.scan_bytes_leftmost_from_all_overlaps only as a measurement path. It first runs the raw overlapping scan, then uses the existing entity-independent shards to reconstruct the exact leftmost-first output. This keeps raw overlap cost and exact reconstruction cost visible without pretending that raw candidates alone preserve enough ordering information. Reconstruction is exact only when the raw overlapping scan itself fits the MatchBuffer pre-scan capacity cap; extremely dense raw overlap workloads can fail before the reconstruction pass runs.

source = b"""
{"entity":"PERSON","canonical_name":"Sam","surface_name":"Sam","regex":"Sam","priority":0}
{"entity":"PERSON","canonical_name":"Samwise","surface_name":"Samwise","regex":"Samwise","priority":1}
{"entity":"PROJECT","canonical_name":"Samba","surface_name":"Samba","regex":"Samba","priority":0}
"""
default_bank = _engine.Bank.from_source_bytes(source, format_hint="jsonl")
overlap_bank = _engine.Bank.from_source_bytes(
    source,
    format_hint="jsonl",
    compile_options_json='{"match_mode":"all_overlaps"}',
)

raw = overlap_bank.scan_bytes(b"Samba Samwise")
default_raw = default_bank.scan_bytes(b"Samba Samwise")
reconstructed = overlap_bank.scan_bytes_leftmost_from_all_overlaps(b"Samba Samwise")

assert [raw[i] for i in range(len(raw))] == [
    (0, 0, 3),
    (2, 0, 5),
    (0, 6, 9),
    (1, 6, 13),
]
assert [reconstructed[i] for i in range(len(reconstructed))] == [default_raw[i] for i in range(len(default_raw))]

Global Leftmost Internal Baseline

compile_options_json='{"match_mode":"global_leftmost"}' builds one combined regex-automata meta matcher in LeftmostFirst mode. It is exposed only as an internal throughput baseline. metadata()["match_mode"] labels it with:

{
  "name": "global_leftmost",
  "status": "internal_benchmark_only",
  "production_default": false,
  "internal_only": true,
  "semantic_notes": "collapses cross-entity overlap to one leftmost-first winner per region and is not semantically equivalent to the production default"
}

The mode intentionally violates NERB's production overlap contract:

source = b"""
{"entity":"PERSON","canonical_name":"Sam","surface_name":"Sam","regex":"Sam","priority":0}
{"entity":"PROJECT","canonical_name":"Samba","surface_name":"Samba","regex":"Samba","priority":0}
"""
default_bank = _engine.Bank.from_source_bytes(source, format_hint="jsonl")
global_bank = _engine.Bank.from_source_bytes(
    source,
    format_hint="jsonl",
    compile_options_json='{"match_mode":"global_leftmost"}',
)

default_raw = default_bank.scan_bytes(b"Samba ships")
global_raw = global_bank.scan_bytes(b"Samba ships")

assert [default_raw[i] for i in range(len(default_raw))] == [(0, 0, 3), (1, 0, 5)]
assert [global_raw[i] for i in range(len(global_raw))] == [(0, 0, 3)]

Native _engine.Bank.scan_path reads one explicit file path in Rust, validates the bytes through the same UTF-8 scanner, and returns raw matches in a MatchBuffer. It does not allocate Python match records. The public Python nerb.Bank.scan_path wrapper uses the native path scan variant that returns the scanned byte snapshot with the raw matches, then projects that same snapshot into public records. Path scans use the same 10 MiB ceiling and bound the read before passing bytes to the scanner.

Error Boundary

Native validation and parse failures are translated to ValueError. File-read failures are OSError. Buffer indexing failures are IndexError, native allocation failures are MemoryError, and panic-safe wrappers translate an unexpected Rust panic into RuntimeError instead of unwinding through Python.

Public Python Bank

from nerb import Bank exposes the high-level Rust-backed wrapper. It projects raw native matches into the public record schema:

from nerb import Bank

bank = Bank.from_source_bytes(b'{"ARTIST":{"Rush":"Rush"}}', format_hint="json")
records = bank.scan_text("Café Rush")
assert records == [
    {
        "entity": "ARTIST",
        "canonical_name": "Rush",
        "surface_name": "Rush",
        "string": "Rush",
        "start": 6,
        "end": 10,
        "offset_unit": "byte",
    }
]

Byte offsets are the default for scan_text, scan_bytes, scan_path, CLI extraction, and MCP extraction. Text callers may explicitly ask for character offsets:

assert bank.scan_text("Café Rush", offsets="char")[0]["offset_unit"] == "char"

scan_text and scan_bytes also accept a positive max_matches keyword. The native collector aborts as soon as that count would be exceeded, before Python record projection or sorting; evaluator code uses this boundary for resource limits.

Bank.scan_path(path) reads the exact file bytes and then uses the native UTF-8 scan path. Invalid UTF-8 raises ValueError; callers that need lossy or custom decoding must decode text explicitly and pass it to scan_text.

Bank.from_config(..., word_boundaries=True) passes the boundary policy to Rust canonicalization. Rust emits canonical JSON with defaults.word_boundaries: true, wraps whole detector regexes once during canonicalization, and includes that policy in pattern stable IDs and the bank hash.

CLI nerb extract and the config-backed MCP extraction tools use this wrapper. Their records expose canonical_name and surface_name.

uv run nerb extract --all --text "Rush played rock." \
  --detector "ARTIST:Rush=Rush" \
  --detector "GENRE:Rock=rock" \
  --format json
[
  {"entity": "ARTIST", "canonical_name": "Rush", "surface_name": "Rush", "string": "Rush", "start": 0, "end": 4, "offset_unit": "byte"},
  {"entity": "GENRE", "canonical_name": "Rock", "surface_name": "Rock", "string": "rock", "start": 12, "end": 16, "offset_unit": "byte"}
]

Compiled Bank Cache And Batch Extraction

The public Python wrapper caches compiled native Bank objects in process. The cache key is semantic rather than path-based:

from nerb import Bank, bank_cache_info, clear_bank_cache

clear_bank_cache()
first = Bank.from_config({"ARTIST": {"Rush": "Rush"}})
second = Bank.from_config({"ARTIST": {"Rush": "Rush"}})

assert first.cache_metadata()["hit"] is False
assert second.cache_metadata()["hit"] is True
print(second.cache_metadata()["key"])
print(bank_cache_info())

Example key shape:

{
  "bank_hash": "sha256:...",
  "schema_version": 1,
  "semantic_version": "0.0.10",
  "engine_name": "nerb_engine",
  "engine_version": "0.0.10",
  "canonical_engine": "rust-regex-meta",
  "compile_options": {"match_mode": "entity_independent"},
  "target_triple": "x86_64-linux-gnu",
  "platform": "linux-x86_64",
  "pointer_width": 64,
  "endian": "little"
}

use_cache=False bypasses lookup and insertion for callers that need isolated compilation. clear_bank_cache() clears only this process. The process-local cache uses bounded LRU eviction and reports max_entries plus max_source_keys in bank_cache_info(). The cache does not serialize matcher state, write engine artifacts, or add a disk cache.

The config-backed MCP extraction tools return the same per-extraction cache metadata and expose engine_cache_info plus clear_engine_cache for process-local diagnostics.

Batch CLI extraction compiles once and scans many explicit documents:

uv run nerb extract-batch doc-a.txt doc-b.txt --entity ARTIST --config detectors.yaml --format json
uv run nerb extract-batch --manifest docs.txt --all --config detectors.yaml --format jsonl
uv run nerb extract-batch --stdin --entity ARTIST --config detectors.yaml --format table

The JSON output includes top-level cache metadata and document payloads in input order. Manifest files are UTF-8 text files with one explicit path per nonblank line; relative paths resolve against the manifest file's parent directory. Recursive walking, gitignore discovery, Rayon batch parallelism, and serialized DFA or engine-payload caches are not part of the current process-local cache.