PlainLogic
← Back to the lab

Interactive lab · Data engineering

Data Pipeline Lab

Twelve messy signup records go in one end; a trustworthy dashboard comes out the other. Watch each record flow through source, ingest, validation, cleaning, transformation, and storage — then break the pipeline on purpose and fix it.

Simulated dataset of twelve made-up records. Everything runs in your browser — nothing is uploaded anywhere.

The simulator

Messy signups in, dashboard out

This is ETL — extract, transform, load — with the two steps real pipelines spend most of their time on: cleaning and validation. Press Run pipeline and follow the dots. Each dot is one signup record; watch what each stage does to it.

Mode

Repair challenge: the CLEAN step's “normalize email” rule is off, so VALIDATE will reject the messy emails. Run the pipeline, read the rejected-records log, then click the CLEAN stage to find the switch.

record in flightrejectedClick any stage to inspect its rules. Records move through concurrently, like an assembly line.

See the raw dataset (12 records)
#nameemailzipsignup_date
101Ava Martinezava.m@example.com339012026-08-12
102John Reyes" JOHN@EMAIL.COM "339022026-08-19
103Sofia Patel(missing)339032026-08-21
104Noah Kimnoah.kim@example.comABCDE2026-08-25
105Emma Liuemma.liu@example.com339042026-13-45
106Ava Martinezava.m@example.com (duplicate of #101)339012026-08-12
107" olivia brown "olivia.b@sample.org339092026-09-02
108Ethan WrightETHAN.W@MAIL.IO339122026-09-15
109Mia Garciamia.garcia@example339132026-09-18
110Lucas Moorelucas.m@example.com33902026-09-21
111Isabella Rossisabella.r@test.net3399009/24/2026
112James Parkjames.p@example.com339142026-09-29

Made-up data. Messy cells are highlighted; quotes show where stray spaces hide.

Metrics

Pipeline health

Inputs12
Accepted–
Rejected–
Runtime (simulated)–
Current stageidle

Throughput per stage

Records entering each stage. The bars shrink wherever a stage rejects records — that shrinkage is the pipeline doing its job.

Inspector

Select a stage

Click any stage above to see what it does and the exact rules it runs. Tip: the repair challenge involves the CLEAN stage.

Event log

What happened

  1. readyChoose a mode and press Run pipeline. The dataset and timings are simulated.

Rejected records

With reasons

  1. No rejections yet — run the pipeline.

Three jobs

Validate, clean, transform — not the same thing

Every data pipeline does these three jobs in the middle, and mixing them up is where most data bugs come from.

Judges · changes nothing

Validation

The bouncer. It checks each record against rules — a zip must be five digits, a date must be a real calendar day — and never edits the data. Record #104's zip ABCDE fails validation, so the record is rejected with a reason. Validation answers one question: is this record acceptable as it stands?

Fixes formatting · never facts

Cleaning

The tidy-up. It repairs formatting, not substance: " JOHN@EMAIL.COM " becomes "john@email.com" via trim + lowercase. Cleaning makes good data consistent — but it cannot rescue a missing email or turn letters into a zip code. If cleaning can't fix it, the record moves on to be judged.

Reshapes for analysis

Transformation

The reshape. It derives new fields the dashboard needs: the email domain for a signups-by-domain chart, the signup month for a trend. Transformation assumes the data is already clean and valid — which is exactly why it runs after both.

A design decision

Why reject instead of guess?

Every rejected record in the log is a decision the pipeline refused to make on its own. Take record #104: zip ABCDE. No cleaning rule can turn letters into a zip code — and if the pipeline invented 33901 to keep things moving, every report downstream would be quietly, confidently wrong.

A guess that looks like data is worse than a rejection you can see. The rejected record lands in a log with a plain-language reason, a human fixes the source system, and the next run passes cleanly. That loop — reject loudly, fix at the source — is what makes a dashboard trustworthy: everything on it survived the same rules, and everything that didn't is listed with a reason.

Notice the pipeline never deletes anything silently, either. The duplicate record #106 isn't merged or overwritten — it's dropped at the database with a note saying which record was kept. Loud beats clever, every time.

Try it

Fix the broken pipeline in four steps

  1. Run the broken pipeline.

    Press Run pipeline in challenge mode. Watch the dots flow left to right — and watch VALIDATE reject the messy emails while the rejected-records log explains why.

  2. Open the CLEAN stage.

    Click the CLEAN node. Its rule list shows “Normalize email (trim + lowercase)” switched off, labeled runs before validation. This pipeline is configured to clean before it validates: formatting fixes run first, then validation judges what remains.

  3. Enable the rule and rerun.

    Flip the switch, press Run pipeline again. The two messy emails are normalized before validation — " JOHN@EMAIL.COM " becomes "john@email.com" — and sail through. Nothing else changes: the genuinely broken records are still rejected, correctly.

  4. Try free run.

    Switch to free-run mode to watch a healthy pipeline end to end, or click any stage to inspect exactly what it checks. Every rule in the inspector is a real function — toggling one genuinely changes the run.

Behind the build

What powers this experiment

ETL, honestly
Extract, clean, validate, transform, load — the unglamorous middle of every data project. The lab teaches the real distinction: cleaning fixes formatting, validation judges truth, and transformation reshapes for analysis.
Deterministic simulation
A fixed twelve-record dataset and pure rule functions: the same input with the same rule switches always produces the same output. There is no randomness anywhere in a run.
Pipeline parallelism
Records advance one stage per beat, all in flight at once — like an assembly line, which is how real pipelines keep throughput up. Counts, metrics, and charts update live as the dots move.
Hand-rolled SVG
The throughput sparkline and the dashboard charts are plain SVG drawn by the page's own script. No chart library — a dozen bars don't need one.
Rules are data
Each stage's rules are real functions, not decoration. Flipping the “normalize email” switch genuinely changes what VALIDATE sees, which is why the repair challenge actually works.

Questions

Fair questions

What is ETL?

ETL stands for Extract, Transform, Load: pulling data out of where it is born, reshaping it, and loading it where it will be used. This lab shows the fuller modern version — extract, clean, validate, transform, load — because real pipelines spend most of their effort on the cleaning and checking in the middle.

What is the difference between validation and cleaning?

Cleaning fixes formatting: trimming spaces, lowercasing an email, reformatting a date. Validation checks truth: is this a real calendar date, is the zip five digits, is the email actually present? Cleaning changes how a value looks; validation judges whether the record is acceptable. A record that cleaning cannot fix is rejected, never guessed.

Why do pipelines reject records instead of guessing?

A guess that looks like data is worse than a rejection you can see. If the pipeline invented a zip code for “ABCDE”, every report downstream would be quietly wrong. A rejected record lands in a log with a plain-language reason, so a human can fix the source. Rejection is a feature, not a failure.

Is my data uploaded anywhere?

No. The dataset is twelve made-up signup records, and the entire pipeline runs in your browser tab. Nothing you do here leaves your device.

Is it free?

Yes. Like everything in the PlainLogic lab, the Data Pipeline Lab is free to use as often as you like. No account, no sign-up, no paid tier.

Keep exploring

More from the lab

Live data

Situation Room

Real pipelines carry real data. Watch live weather, alerts, earthquakes, and radar flow through an operations console.

Interactive lab

Network Operations Lab

Another simulator from the bench: send packets through switches and routers, inject faults, and learn to diagnose them.

The lab

PlainLogic home

Games, experiments, tools, and the blog — the full bench.