Named Entity Regex Builder¶
NERB is a Python package, CLI, and MCP server for curated named-entity banks. Define known entities once, validate them before use, scan text locally with the Rust-backed engine, and return byte-offset JSON records that agents and services can cite, patch, diff, evaluate, and promote.
pip install --upgrade nerb
nerb validate-bank --bank company.json
nerb extract-text \
--bank company.json \
--text "Send this to Acme Corp today."
Quickstart Schema Performance Enron charter Prepare Enron Split Enron Build an Enron bank Evaluate Enron Measure Enron performance Enron evidence and decision
Compile knowledge once, scan many times¶
An entity bank is a reviewable cache of approved names, aliases, and structured patterns. NERB deterministically finds qualifying cataloged occurrences under the bank's declared matching semantics. Entities absent from the bank are outside that guarantee; applications that need open-ended discovery must measure bank coverage or add a separately evaluated discovery layer.
Enron evidence: known-bank contract passed
NERB detected and correctly mapped 39,604/39,604 approved positive cases across all 13,201 active patterns, with 1,210/1,210 required negative cases clean. A separate natural-text diagnostic found 142/146 cataloged exact matches. The constructed bank knew only 146/1,393 labeled spans, so it is not suitable by itself as a comprehensive PII redactor. Read the evidence interpretation for the guarantee, coverage, and scale results.
Why Teams Use NERB¶
Known Entities¶
Use curated names, aliases, domains, codes, accounts, products, vendors, people, or compliance terms when open-domain NER is not the right control surface.
Deterministic Records¶
Get stable entity IDs, canonical names, matched strings, and byte offsets for evidence-backed reports, redaction, diffs, evals, and CI gates.
One Local Surface¶
Run the same bank through Python helpers, shell commands, and local MCP tools without sending documents to a hosted model or rewriting extraction logic.
Complete Core Loop¶
Create a minimal JSON bank:
{
"schema_version": "nerb.bank.v1",
"id": "company_entities",
"name": "Company Entities",
"description": "Companies to recognize in internal documents.",
"version": "2026.06.24",
"status": "active",
"created_at": "2026-06-24T00:00:00Z",
"updated_at": "2026-06-24T00:00:00Z",
"unicode_normalization": "none",
"default_regex_flags": ["IGNORECASE"],
"entities": {
"company": {
"description": "Organizations.",
"status": "active",
"regex_flags": [],
"names": {
"acme_corp": {
"canonical": "Acme Corp",
"description": "Primary account.",
"status": "active",
"patterns": {
"primary": {
"kind": "literal",
"value": "Acme Corp",
"description": "Exact company alias.",
"status": "active",
"priority": 100,
"case_sensitive": false,
"normalize_whitespace": true,
"left_boundary": "word",
"right_boundary": "word",
"metadata": {}
}
},
"metadata": {}
}
},
"metadata": {}
}
},
"metadata": {}
}
Validate and scan:
nerb validate-bank --bank company.json
nerb extract-text \
--bank company.json \
--text "Send this to Acme Corp today."
NERB returns deterministic JSON records:
{
"records": [
{
"entity": "company",
"canonical_name": "Acme Corp",
"surface_name": "Acme Corp",
"string": "Acme Corp",
"start": 13,
"end": 22,
"offset_unit": "byte",
"entity_id": "company",
"name_id": "acme_corp",
"pattern_id": "primary",
"pattern_kind": "literal",
"captures": {}
}
]
}
Choose Your Path¶
- Quickstart
- Install NERB, validate a bank, extract from text, and call the Python helper.
- Workflows
- Build, validate, patch, diff, evaluate, benchmark, and regress entity banks.
- Interfaces
- Map the same extraction behavior across CLI commands, Python helpers, and MCP tools.
- Anonymization
- Replace entities with stable redaction tokens or pseudonyms and restore reversible DBs intentionally.
- Schema Reference
- Read the JSON bank, extraction record, eval, replacement DB, and diagnostic contracts.
- Performance
- Reproduce Rust-backed gates and run the private Enron compile-once/cache-value workflow.
When NERB Fits¶
| Use NERB when you need | Prefer another tool when you need |
|---|---|
| Known, curated entities with reviewable aliases | Open-domain entity discovery |
| Local processing for sensitive documents | Hosted extraction or human annotation workflows |
| Stable byte offsets and source IDs | Probabilistic labels without record contracts |
| CI gates for bank changes | One-off exploratory extraction with no promotion path |
Performance Evidence¶
| Workload | Patterns | Scan/project median | Throughput |
|---|---|---|---|
| Medium production bank | 8,000 | 0.008654s | 11.6 MB/s |
| 1 MB evidence run | 8,000 | 0.043692s | 22.9 MB/s |
The scale chart compares three synthetic banks: 1,000 patterns over 49,983 document bytes at 10.68 MB/s, 4,000 patterns over 149,995 bytes at 8.59 MB/s, and 10,000 patterns over 299,991 bytes at 6.15 MB/s. The corresponding record counts were 1,136, 3,409, and 6,818.
Reproduce the gate with:
uv run python scripts/rust_engine_gate_report.py --iterations 5 --target-bytes 100000 --dense-bytes 512 \
--bank-owner-entity-count 1000 \
--bank-owner-growth-entity-count 1000 \
--bank-owner-note "representative synthetic medium bank target"
See Performance And Scale Evidence for the full report context.