How we test
The rules we hold ourselves to, written down before the first result so we can't bend them to fit one.
Last updated October 6, 2026
The method is still a draft. We'll pilot it on all three machines, fix whatever turns out to be ambiguous, and only then freeze version 1. Any result published before that will say which draft produced it.
Three ways we run a printer
Stock baseline. The manufacturer's own slicer profile, with only the shared requirements changed: nozzle size, layer height, filament, model, and which color goes on which tool. We record every change we had to make.
Tuned. Our best effort on each machine, with the tuning and the time it took written down. Tuned results never sit beside another printer's baseline without a clear warning.
Modified. Different firmware or hardware is a different configuration. We capture a stock baseline before changing anything, so you can see the before and after.
What every run records
- The machine, its configuration, firmware, and slicer version, plus the exported project file.
- Each filament: brand, color, how it was dried, and which tool it was loaded on.
- The slicer's time estimate, the actual start and end, and the outcome.
- Requested and observed tool changes, retries, and every time we had to intervene.
- Separate weights for the object, supports and brim, prime tower, purge, and any failed attempts.
- Photos of the object and of each waste pile, taken the same way every time.
A run ends one of four ways: completed, completed with our help, failed, or aborted. Failures are published with the reason and the stage they happened at.
Time and tool changes
"How long did it take" needs a start and an end. We time from starting the job to the printer reporting it finished, and log preheat and calibration separately where the machine lets us.
A tool change has two different durations. The mechanical swap is the head leaving and the next one locking in. The real cost to your print also includes travel, waiting for temperature, wiping, and priming. Manufacturer figures usually describe the first; we report both and say exactly where each one starts and stops.
A thousand tool changes in one print are a thousand observations from one job, not a thousand independent tests. We report completed jobs, observed changes, failed changes, and interventions separately. "No failures in these runs" is something we can say. "100% reliable" isn't.
Waste
Supports you'd need on any printer are a different thing from material burned by switching colors, so we weigh them separately. When we give an overall figure, it's the weight of everything that isn't the object, divided by the weight of the object plus everything we captured. If some purge escaped the bin, we'll say so.
Benchmarks
Candidate tests. None are frozen yet.
| Test | The question it answers | Status |
|---|---|---|
| TCB-001Four-tool reference print | What does a comparable everyday multi-color print look like, and what does it cost in time and material? | Candidate |
| TCB-002High-swap endurance | What happens over a large number of real tool changes? | Candidate |
| TCB-003Rigid plus flexible | Can the workflow produce a useful multi-material part? | Candidate |
| TCB-004Support interfaces | How cleanly can difficult supported surfaces be produced? | Candidate |
| TCB-005Tool registration | How well do features printed by different tools line up? | Candidate |
| TCB-006Firmware and modification retest | What changed after an update or modification? | Candidate |
What our results can't tell you
A result describes one machine, in one configuration, on one firmware, printing one workload. A later update can change it, and when we retest, the old result stays up next to the new one.
We aim for three runs of each baseline per machine, and we'll publish however many we actually managed. A single run is one observation and is labeled preliminary. We don't give overall scores or star ratings.
How each number is labeled
- Manufacturer specification Taken from an official source, linked with the date we checked it.
- Our measurement Measured by us and linked to its run.
- Community report Reported by other owners and not reproduced by us.
- Not tested We don't know yet.