Rust Engine PyO3 Boundary¶
This document records the native boundary that backs the current Python Bank wrapper and the CLI/MCP extraction
surfaces, including the production matcher and the two internal measurement modes.
Native Bank Constructors¶
from nerb import _engine
bank = _engine.Bank.from_source_bytes(b'{"CODE":{"Alpha":"A"}}', format_hint="json")
canonical = bank.to_canonical_json_bytes()
round_tripped = _engine.Bank.from_canonical_json_bytes(canonical)
assert round_tripped.metadata()["bank_hash"] == bank.metadata()["bank_hash"]
metadata() exposes only engine and canonical-bank facts needed by Python wrappers:
{
"engine": "nerb_engine",
"build_source_sha256": "sha256:...",
"schema": 1,
"bank_hash": "sha256:...",
"entity_count": 1,
"pattern_count": 1,
"defaults": {
"engine": "rust-regex-meta",
"unicode": true,
"case_insensitive": false,
"word_boundaries": false,
"normalization": "none"
},
"compile_options": {
"match_mode": "entity_independent"
},
"match_mode": {
"name": "entity_independent",
"status": "production_default",
"production_default": true,
"internal_only": false,
"semantic_notes": "reports cross-entity overlap with leftmost-first matching within each entity"
},
"scan_limits": {
"maximum_input_bytes": 10485760,
"maximum_concurrent_scans_per_bank": 8
},
"regex_resources": {
"scope": "entity_independent_shards",
"physical_regex_layers": 0,
"maximum_regex_layers_per_entity": 0,
"compiled_regex_static_bytes": 0,
"eager_cache_bytes_per_scan": 0,
"pikevm_cache_projection_bytes_per_scan": 0,
"pikevm_stack_growth_allowance_bytes_per_scan": 0,
"lazy_dfa_growth_allowance_bytes_per_scan": 0,
"regex_cache_allowance_bytes": 0,
"size_limit_bisections": 0,
"resource_limit_bisections": 0,
"accounted_bytes": 0,
"cache_concurrency_budget": 8,
"explicit_regex_cache_slots": 8,
"internal_meta_cache_pool_used": false,
"per_lazy_dfa_cache_capacity_bytes": 32768,
"maximum_lazy_dfa_caches_per_regex": 3,
"pikevm_stack_nfa_memory_multiplier": 16,
"onepass_enabled": false,
"bounded_backtracker_enabled": false,
"maximum_layers_per_entity": 128,
"maximum_patterns_per_regex_layer": 128,
"maximum_accounted_bytes": 805306368
},
"detectors": [
{
"detector_index": 0,
"entity": "CODE",
"canonical_name": "Alpha",
"surface_name": "Alpha",
"stable_id": "pattern:sha256:...",
"priority": 0
}
]
}
The module-level _engine.BUILD_SOURCE_SHA256 equals each bank's build_source_sha256. The build hashes a closed,
sorted inventory containing Cargo.toml, Cargo.lock, build.rs, and every Rust source file after LF normalization;
the build fails if that inventory changes without an explicit update. Only production entity_independent metadata
contains regex_resources, because its accounting does not describe the internal modes' additional matchers.
MatchBuffer¶
MatchBuffer is a Rust-owned container for raw scan results:
Python can create, reserve, clear, and inspect a buffer without constructing public record dictionaries:
buffer = _engine.MatchBuffer(capacity=1024)
assert len(buffer) == 0
raw = _engine.MatchBuffer.from_raw_matches([(7, 10, 15)])
assert raw[0] == (7, 10, 15)
raw.clear()
from_raw_matches accepts a sized Python sequence and exists to test the boundary. Bank.scan_bytes fills
MatchBuffer from Rust. Public record projection remains outside the scan loop.
Python-created buffers and Rust scanner appends are capped at 1,000,000 requested raw matches and use fallible Rust allocation paths. Later dense-hit measurement may revisit this logical limit.
Scanning¶
Bank.scan_bytes implements the production-default entity_independent mode: one logical matcher per entity with
leftmost-first semantics inside each entity and cross-entity overlap preserved. A logical matcher may contain bounded
exact-literal Aho-Corasick layers, mapped Aho-Corasick layers for supported normalized-whitespace and simple-fold
literals, and one or more residual regex-automata layers. Global arbitration by start offset and original pattern order
reconstructs the entity's leftmost-first result. It validates UTF-8 input,
releases the GIL during the Rust scan, returns raw (detector_index, start_byte, end_byte) matches sorted by byte offsets
and detector index, and optionally fills a caller-provided MatchBuffer.
Matcher construction disables one-pass and bounded-backtracker strategies and applies bounded regex-automata NFA,
hybrid-cache, and DFA limits. Size-limit failures
for advancing residual patterns are deterministically bisected, with explicit per-entity layer and aggregate
static/cache-memory ceilings. Unsupported syntax, singleton size failures, or aggregate resource exhaustion fail during
bank construction before the bank is returned; syntax/shape failures raise ValueError, while aggregate memory-budget
exhaustion raises MemoryError. Production metadata reports the realized layer, memory, and bisection profile.
Cache accounting includes every physical regex, including the first or only regex in an entity. NERB does not use the
meta regex's internal sharded cache pool: each physical regex owns exactly eight explicit cache slots selected by the
bank's scan permit. For one scan slot, the allowance sums the eager meta-cache heap, a checked projection of the
equivalent PikeVM fallback's fixed cache, a conservative PikeVM epsilon-stack growth bound, and three 32 KiB lazy-DFA
capacities (forward, reverse, and reverse-inner). regex_cache_allowance_bytes multiplies that per-scan sum by eight, and
accounted_bytes adds compiled static bytes. One-pass and bounded-backtracker strategies are disabled so no unmeasured
lazy strategy cache can appear.
An initial physical regex layer contains at most 128 patterns. This named envelope bounds the PikeVM's implicit-capture state/slot product before compilation while retaining deterministic pattern order. A layer that still exceeds a compile-size or accounted-resource limit is bisected deterministically; metadata reports size-limit and resource-limit bisections separately. More than 128 physical layers in one entity or more than 768 MiB of aggregate static-plus-eight-slot cache allowance fails before the layer is committed.
The PikeVM stack bound is 16 times the compiled implicit-capture NFA's reported memory. In pinned regex-automata
0.4.14, every epsilon-stack push is backed by an NFA state or stored Union alternate; the largest private frame is under
four machine words. The multiplier covers the worst frame-to-StateID ratio and geometric Vec retained capacity.
The fixed-cache projection uses checked arithmetic for the four NFA-state ID vectors and two capture-slot tables that
regex-automata creates. Production accounting therefore rejects arithmetic overflow without first allocating a
potentially quadratic table. Safe-size regressions compare the projection with an actual PikeVM cache, including a
complex cache larger than 32 KiB, and exercise overflow failure directly.
Every compiled bank admits at most eight scans at once. A per-bank permit covers every native scan mode and is released on success, validation/allocation failure, or panic unwinding; additional callers wait without reducing the first eight to a serial lane. This enforced scan ceiling is the same concurrency value used by regex-cache accounting.
All inline scan variants reject inputs larger than 10 MiB before releasing the GIL or allocating mapped-haystack
projections. Exactly 10 MiB is accepted. scan_text inherits this byte limit after UTF-8 encoding, so a Unicode string's
encoded length—not its Python character count—is authoritative. Reused MatchBuffer objects are cleared on an
over-limit error, just as they are for other scan failures.
IGNORECASE, MULTILINE, DOTALL, and VERBOSE are applied through per-pattern syntax configuration. The ASCII flag
lowers ASCII-sensitive escapes and boundaries such as \w, \d, \s, and \b while leaving the rest of the detector
pattern in UTF-8-safe Unicode regex mode.
bank = _engine.Bank.from_source_bytes(b'{"PERSON":{"Sam":"Sam"},"PROJECT":{"Samba":"Samba"}}')
raw = bank.scan_bytes(b"Samba ships")
assert [raw[i] for i in range(len(raw))] == [(0, 0, 3), (1, 0, 5)]
All Overlaps Prototype¶
compile_options_json='{"match_mode":"all_overlaps"}' builds an internal prototype around lower-level
regex-automata hybrid DFAs:
- a forward DFA runs overlapping search with
MatchKind::All; - a reverse DFA with per-pattern start states recovers the start byte for each reported end;
- local pattern IDs are translated back to global detector indexes before appending to
MatchBuffer.
The prototype rejects Unicode word-boundary assertions such as \b because the lower-level DFA only provides heuristic
Unicode-boundary support that can quit on valid non-ASCII UTF-8. Use explicit ASCII word-boundary syntax such as
(?-u:\b) for raw all_overlaps, or use the production-default entity_independent mode for Unicode boundary
semantics.
Raw all_overlaps output is intentionally not the default contract. It preserves cross-entity overlap, but it also
reports within-entity overlapping detectors and every matching span for each detector pattern. It does not preserve a
separate branch identity inside one regex; attribution still stops at the NERB detector index. For example, the
production entity_independent mode chooses Samwise for Samwise|Sam over Samwise, while raw all_overlaps exposes
both (0, 0, 3) and (0, 0, 7) for that one detector. That means a span-only candidate post-filter cannot prove exact
leftmost-first reconstruction.
The prototype therefore exposes Bank.scan_bytes_leftmost_from_all_overlaps only as a measurement path. It first runs
the raw overlapping scan, then uses the existing entity-independent shards to reconstruct the exact leftmost-first output.
This keeps raw overlap cost and exact reconstruction cost visible without pretending that raw candidates alone preserve
enough ordering information. Reconstruction is exact only when the raw overlapping scan itself fits the MatchBuffer
pre-scan capacity cap; extremely dense raw overlap workloads can fail before the reconstruction pass runs.
source = b"""
{"entity":"PERSON","canonical_name":"Sam","surface_name":"Sam","regex":"Sam","priority":0}
{"entity":"PERSON","canonical_name":"Samwise","surface_name":"Samwise","regex":"Samwise","priority":1}
{"entity":"PROJECT","canonical_name":"Samba","surface_name":"Samba","regex":"Samba","priority":0}
"""
default_bank = _engine.Bank.from_source_bytes(source, format_hint="jsonl")
overlap_bank = _engine.Bank.from_source_bytes(
source,
format_hint="jsonl",
compile_options_json='{"match_mode":"all_overlaps"}',
)
raw = overlap_bank.scan_bytes(b"Samba Samwise")
default_raw = default_bank.scan_bytes(b"Samba Samwise")
reconstructed = overlap_bank.scan_bytes_leftmost_from_all_overlaps(b"Samba Samwise")
assert [raw[i] for i in range(len(raw))] == [
(0, 0, 3),
(2, 0, 5),
(0, 6, 9),
(1, 6, 13),
]
assert [reconstructed[i] for i in range(len(reconstructed))] == [default_raw[i] for i in range(len(default_raw))]
Global Leftmost Internal Baseline¶
compile_options_json='{"match_mode":"global_leftmost"}' builds one combined regex-automata meta matcher in
LeftmostFirst mode. It is exposed only as an internal throughput baseline. metadata()["match_mode"] labels it with:
{
"name": "global_leftmost",
"status": "internal_benchmark_only",
"production_default": false,
"internal_only": true,
"semantic_notes": "collapses cross-entity overlap to one leftmost-first winner per region and is not semantically equivalent to the production default"
}
The mode intentionally violates NERB's production overlap contract:
source = b"""
{"entity":"PERSON","canonical_name":"Sam","surface_name":"Sam","regex":"Sam","priority":0}
{"entity":"PROJECT","canonical_name":"Samba","surface_name":"Samba","regex":"Samba","priority":0}
"""
default_bank = _engine.Bank.from_source_bytes(source, format_hint="jsonl")
global_bank = _engine.Bank.from_source_bytes(
source,
format_hint="jsonl",
compile_options_json='{"match_mode":"global_leftmost"}',
)
default_raw = default_bank.scan_bytes(b"Samba ships")
global_raw = global_bank.scan_bytes(b"Samba ships")
assert [default_raw[i] for i in range(len(default_raw))] == [(0, 0, 3), (1, 0, 5)]
assert [global_raw[i] for i in range(len(global_raw))] == [(0, 0, 3)]
Native _engine.Bank.scan_path reads one explicit file path in Rust, validates the bytes through the same UTF-8 scanner,
and returns raw matches in a MatchBuffer. It does not allocate Python match records. The public Python
nerb.Bank.scan_path wrapper uses the native path scan variant that returns the scanned byte snapshot with the raw
matches, then projects that same snapshot into public records. Path scans use the same 10 MiB ceiling and bound the read
before passing bytes to the scanner.
Error Boundary¶
Native validation and parse failures are translated to ValueError. File-read failures are OSError. Buffer indexing
failures are IndexError, native allocation failures are MemoryError, and panic-safe wrappers
translate an unexpected Rust panic into RuntimeError instead of unwinding through Python.
Public Python Bank¶
from nerb import Bank exposes the high-level Rust-backed wrapper. It projects raw native matches into the public record
schema:
from nerb import Bank
bank = Bank.from_source_bytes(b'{"ARTIST":{"Rush":"Rush"}}', format_hint="json")
records = bank.scan_text("Café Rush")
assert records == [
{
"entity": "ARTIST",
"canonical_name": "Rush",
"surface_name": "Rush",
"string": "Rush",
"start": 6,
"end": 10,
"offset_unit": "byte",
}
]
Byte offsets are the default for scan_text, scan_bytes, scan_path, CLI extraction, and MCP extraction. Text callers
may explicitly ask for character offsets:
scan_text and scan_bytes also accept a positive max_matches keyword. The native collector aborts as soon as that
count would be exceeded, before Python record projection or sorting; evaluator code uses this boundary for resource
limits.
Bank.scan_path(path) reads the exact file bytes and then uses the native UTF-8 scan path. Invalid UTF-8 raises
ValueError; callers that need lossy or custom decoding must decode text explicitly and pass it to scan_text.
Bank.from_config(..., word_boundaries=True) passes the boundary policy to Rust canonicalization. Rust emits canonical
JSON with defaults.word_boundaries: true, wraps whole detector regexes once during canonicalization, and includes that
policy in pattern stable IDs and the bank hash.
CLI nerb extract and the config-backed MCP extraction tools use this wrapper. Their records expose
canonical_name and surface_name.
uv run nerb extract --all --text "Rush played rock." \
--detector "ARTIST:Rush=Rush" \
--detector "GENRE:Rock=rock" \
--format json
[
{"entity": "ARTIST", "canonical_name": "Rush", "surface_name": "Rush", "string": "Rush", "start": 0, "end": 4, "offset_unit": "byte"},
{"entity": "GENRE", "canonical_name": "Rock", "surface_name": "Rock", "string": "rock", "start": 12, "end": 16, "offset_unit": "byte"}
]
Compiled Bank Cache And Batch Extraction¶
The public Python wrapper caches compiled native Bank objects in process. The cache key is semantic rather than
path-based:
from nerb import Bank, bank_cache_info, clear_bank_cache
clear_bank_cache()
first = Bank.from_config({"ARTIST": {"Rush": "Rush"}})
second = Bank.from_config({"ARTIST": {"Rush": "Rush"}})
assert first.cache_metadata()["hit"] is False
assert second.cache_metadata()["hit"] is True
print(second.cache_metadata()["key"])
print(bank_cache_info())
Example key shape:
{
"bank_hash": "sha256:...",
"schema_version": 1,
"semantic_version": "0.0.10",
"engine_name": "nerb_engine",
"engine_version": "0.0.10",
"canonical_engine": "rust-regex-meta",
"compile_options": {"match_mode": "entity_independent"},
"target_triple": "x86_64-linux-gnu",
"platform": "linux-x86_64",
"pointer_width": 64,
"endian": "little"
}
use_cache=False bypasses lookup and insertion for callers that need isolated compilation. clear_bank_cache() clears
only this process. The process-local cache uses bounded LRU eviction and reports max_entries plus max_source_keys in
bank_cache_info(). The cache does not serialize matcher state, write engine artifacts, or add a disk cache.
The config-backed MCP extraction tools return the same per-extraction cache metadata and expose engine_cache_info plus
clear_engine_cache for process-local diagnostics.
Batch CLI extraction compiles once and scans many explicit documents:
uv run nerb extract-batch doc-a.txt doc-b.txt --entity ARTIST --config detectors.yaml --format json
uv run nerb extract-batch --manifest docs.txt --all --config detectors.yaml --format jsonl
uv run nerb extract-batch --stdin --entity ARTIST --config detectors.yaml --format table
The JSON output includes top-level cache metadata and document payloads in input order. Manifest files are UTF-8 text
files with one explicit path per nonblank line; relative paths resolve against the manifest file's parent directory.
Recursive walking, gitignore discovery, Rayon batch parallelism, and serialized DFA or engine-payload caches are not part
of the current process-local cache.