scrubbed

One native binary for turning raw web pages into clean training and evaluation data.

Why

Encoding damage, boilerplate, personal information (PII), and duplicate pages quietly degrade training data. Fixing them often means chaining several Python tools. Their dependencies may conflict and force separate environments. Deployment can also involve building or pulling large container images before the first document is processed.

scrubbed puts encoding repair, main-content extraction, and PII scanning in one native binary, with language ID and near-duplicate decisions as optional stages. That means one artifact to deploy and fewer handoffs between tools.

Where this fits in a training pipeline

scrubbed collects pages, cleans files to text, and transforms selected fields in existing JSONL streams.

Collect
scrubbed crawl
→
Clean & curate
scrubbed
→
Package
shards, tokenizer
→
Train
your training run

scrubbed crawl collects raw pages locally; clean-web-document cleans them in a separate pass. scrubbed does not tokenize or train.

See it work

Real output from the actual binary, not a mockup.

Input (UTF-8 bytes misread as Windows-1252)
The café menu featured crème brûlée — a classic dessert.
./scrubbed run --filters fix-mojibake
The café menu featured crème brûlée — a classic dessert.
Input (raw HTML: nav, ad sidebar, footer)
<nav>Home | Archive | Subscribe</nav>
<aside>Sponsored: buy our
  newsletter subscription today!</aside>
<article>
  <h1>Field notes from the
  delta survey</h1>
  <p>The survey team spent three
  weeks mapping the river delta's
  shifting sandbars, recording
  water depth every two hundred
  meters along six transects. The
  main channel has migrated nearly
  forty meters east since the last
  survey.</p>
</article>
<footer>&copy; 2026 Delta Survey
  Project. All rights reserved.</footer>
./scrubbed extract --format main-content-markdown
# Field notes from the delta survey

The survey team spent three weeks mapping the river delta's shifting sandbars, recording water depth every two hundred meters along six transects. The main channel has migrated nearly forty meters east since the last survey.

The pipeline, one command

scrubbed clean-web-document --input page.html --output page.txt
scrubbed clean-web-document --input pages/ --output clean/ --threads 4

Repairs encoding damage, extracts the article, captures metadata, scans for common PII, and writes an audit sidecar.

Benchmarks and comparison checks

Methods: ftfy correctness, language and PII, matched pipeline and extraction, and mojibake speed.

Overall matched pipeline

~37× Matched CPU time vs. warm ftfy → Trafilatura → langdetect → Presidio

On an Apple M4, scrubbed used 1.65 CPU-seconds and the warm Python pipeline used 61.12 over 400 files. This is a workload-specific benchmark, not a universal speed claim. Methodology.

Individual stages

39/39 + 48/48 fix-mojibake vs. ftfy

Exact repair and clean-input agreement within scrubbed's Latin-1/Windows-1252 scope. A separate 5.9 MB fixture ran about 25x faster by whole-process wall time.

89.4% / 76.1% Article-text overlap with Trafilatura

On average across 20 real pages, 89.4% of the words scrubbed kept also appeared in Trafilatura's output; scrubbed kept 76.1% of Trafilatura's words. This measures agreement, not accuracy—neither output is treated as the correct answer.

11/11 language-id-detect vs. langdetect

Classification agreement on fixtures covering 11 of scrubbed's 17 supported languages.

21/21 pii-four-class vs. Presidio

Exact fixture agreement for email, phone, card, and IPv4 detection.

Other capabilities

Examples

CI runs the complete example gallery: quickstart, cleaning, extraction formats, custom stages, near-duplicate decisions, PII policies, metadata routing, and local crawling.

No S3, Parquet, native-Windows, or distributed-execution examples are advertised because those are not shipped CLI capabilities.

Status

Ready for local, single-machine corpus runs. Multi-machine runs under DataTrove, Ray, Spark, Slurm, or Kubernetes have not yet been tested on a representative cluster. Multi-terabyte operation is therefore not yet a proven claim.

Have real distributed infrastructure? We actively want operators to test scrubbed at scale. Open an issue with your orchestrator, corpus size, hardware, throughput, and any failures.

Install

Choose your platform.

Platform

brew install schancel/scrubbed/scrubbed

Homebrew automatically adds the tap. Supports Apple Silicon with macOS 15 (Sequoia) or later.

Need an archive or want to build from source? See the release downloads and quick start.

Usage

Supported: macOS 15+ on Apple Silicon and Linux with glibc 2.36+ on x86_64/aarch64. There is no native Windows build; WSL2 uses the Linux build but has not yet been verified on a Windows host. See the WSL2 evaluation and Linux adapter status for details.

Build

dub build --build=release

Common commands

./scrubbed clean-web-document --input page.html --output page.txt
./scrubbed clean-web-document --input pages/ --output clean/ --threads 4
./scrubbed extract --input page.html --output article.md \
    --format main-content-markdown
./scrubbed crawl --seeds seeds.txt --corpus-dir crawled/ \
    --db crawl.sqlite --concurrency 4 --max-pages 200

Custom pipeline

./scrubbed run --input page.html --output page.out \
    --sidecar-output page.document-metadata.json \
    --stage clean=text-transform --filter fix-mojibake \
    --stage extract=html-main-content \
    --stage scan=pii-four-class \
    --stage publish=document-metadata-publish

JSONL stream

./scrubbed run --input - --output - --jsonl-fields text,title \
    --dataset-namespace corpus-v1 --source-key shard-0001 \
    --max-jsonl-line-bytes 1048576 --max-jsonl-output-bytes 2097152 \
    --filters fix-mojibake < input.jsonl > clean.jsonl

Full command and flag reference: docs/cli-commands.md.

Support

scrubbed is free and MIT-licensed. If it saves you time, you can help fund its maintenance.

Support is optional; the project remains fully open.