scrubbed
One native binary for turning raw web pages into clean training and evaluation data.
Why
Encoding damage, boilerplate, personal information (PII), and duplicate pages quietly degrade training data. Fixing them often means chaining several Python tools. Their dependencies may conflict and force separate environments. Deployment can also involve building or pulling large container images before the first document is processed.
scrubbed puts encoding repair, main-content extraction, and PII scanning in one native binary, with language ID and near-duplicate decisions as optional stages. That means one artifact to deploy and fewer handoffs between tools.
Where this fits in a training pipeline
scrubbed collects pages, cleans files to text, and transforms selected fields in existing JSONL streams.
scrubbed crawl
scrubbed
shards, tokenizer
your training run
scrubbed crawl collects raw pages locally; clean-web-document cleans them in a separate pass. scrubbed does not tokenize or train.
See it work
Real output from the actual binary, not a mockup.
The café menu featured crème brûlée — a classic dessert.
The café menu featured crème brûlée — a classic dessert.
<nav>Home | Archive | Subscribe</nav> <aside>Sponsored: buy our newsletter subscription today!</aside> <article> <h1>Field notes from the delta survey</h1> <p>The survey team spent three weeks mapping the river delta's shifting sandbars, recording water depth every two hundred meters along six transects. The main channel has migrated nearly forty meters east since the last survey.</p> </article> <footer>© 2026 Delta Survey Project. All rights reserved.</footer>
# Field notes from the delta survey The survey team spent three weeks mapping the river delta's shifting sandbars, recording water depth every two hundred meters along six transects. The main channel has migrated nearly forty meters east since the last survey.
The pipeline, one command
scrubbed clean-web-document --input page.html --output page.txt scrubbed clean-web-document --input pages/ --output clean/ --threads 4
Repairs encoding damage, extracts the article, captures metadata, scans for common PII, and writes an audit sidecar.
Benchmarks and comparison checks
Methods: ftfy correctness, language and PII, matched pipeline and extraction, and mojibake speed.
Overall matched pipeline
On an Apple M4, scrubbed used 1.65 CPU-seconds and the warm Python pipeline used 61.12 over 400 files. This is a workload-specific benchmark, not a universal speed claim. Methodology.
Individual stages
Exact repair and clean-input agreement within scrubbed's Latin-1/Windows-1252 scope. A separate 5.9 MB fixture ran about 25x faster by whole-process wall time.
On average across 20 real pages, 89.4% of the words scrubbed kept also appeared in Trafilatura's output; scrubbed kept 76.1% of Trafilatura's words. This measures agreement, not accuracy—neither output is treated as the correct answer.
Classification agreement on fixtures covering 11 of scrubbed's 17 supported languages.
Exact fixture agreement for email, phone, card, and IPv4 detection.
Other capabilities
- Main-content or full-page Markdown export, preserving headings, lists, links, and emphasis.
- PII report, mask, and redact modes for email, phone, card, and IPv4 data.
- Exact-byte deduplication for document shards.
- Deterministic, model-free language identification.
- Quality signals for repetitive or malformed text.
- HTML entity decoding, text normalization, and selective JSONL streaming.
- Parallel file processing, atomic output, and resumable runs.
- A local, multithreaded, resumable crawler.
Examples
CI runs the complete example gallery: quickstart, cleaning, extraction formats, custom stages, near-duplicate decisions, PII policies, metadata routing, and local crawling.
No S3, Parquet, native-Windows, or distributed-execution examples are advertised because those are not shipped CLI capabilities.
Status
Ready for local, single-machine corpus runs. Multi-machine runs under DataTrove, Ray, Spark, Slurm, or Kubernetes have not yet been tested on a representative cluster. Multi-terabyte operation is therefore not yet a proven claim.
Have real distributed infrastructure? We actively want operators to test scrubbed at scale. Open an issue with your orchestrator, corpus size, hardware, throughput, and any failures.
Install
Choose your platform.
Platform
brew install schancel/scrubbed/scrubbed
Homebrew automatically adds the tap. Supports Apple Silicon with macOS 15 (Sequoia) or later.
Need an archive or want to build from source? See the release downloads and quick start.
Usage
Supported: macOS 15+ on Apple Silicon and Linux with glibc 2.36+ on x86_64/aarch64. There is no native Windows build; WSL2 uses the Linux build but has not yet been verified on a Windows host. See the WSL2 evaluation and Linux adapter status for details.
Build
dub build --build=release
Common commands
./scrubbed clean-web-document --input page.html --output page.txt
./scrubbed clean-web-document --input pages/ --output clean/ --threads 4
./scrubbed extract --input page.html --output article.md \
--format main-content-markdown
./scrubbed crawl --seeds seeds.txt --corpus-dir crawled/ \
--db crawl.sqlite --concurrency 4 --max-pages 200
Custom pipeline
./scrubbed run --input page.html --output page.out \
--sidecar-output page.document-metadata.json \
--stage clean=text-transform --filter fix-mojibake \
--stage extract=html-main-content \
--stage scan=pii-four-class \
--stage publish=document-metadata-publish
JSONL stream
./scrubbed run --input - --output - --jsonl-fields text,title \
--dataset-namespace corpus-v1 --source-key shard-0001 \
--max-jsonl-line-bytes 1048576 --max-jsonl-output-bytes 2097152 \
--filters fix-mojibake < input.jsonl > clean.jsonl
Full command and flag reference: docs/cli-commands.md.
Support
scrubbed is free and MIT-licensed. If it saves you time, you can help fund its maintenance.
Support is optional; the project remains fully open.