How to Evaluate an Emerging Treatment Technology: A Field-Tested Framework
A new treatment technology is only as good as the evaluation behind it. A field-tested framework — clear objectives, real conditions, independence, and a path to a design basis — separates a defensible result from a hopeful one.
Most pilots answer the wrong question. They confirm that a technology can work, somewhere, sometime, under conditions the operator was watching closely. What a utility actually needs to know is harder: will this process meet our objectives on our wastewater, every day, for years, run by the staff we already have. A demonstration that produces a single good month of data answers the first question and tells you almost nothing about the second. The gap between them is where a lot of capital gets spent on processes that disappointed once they were built. A disciplined evaluation framework closes that gap. It is the difference between a real result and a hopeful one, and that difference lies mostly in the design of the test, not the cleverness of the technology.
Define success before you test
The single most common evaluation failure is deciding what counts as a good result after the data is in. It is easy, with a season of measurements in hand, to find a metric the technology happened to satisfy and call that the goal. That is not evaluation; it is curve-fitting. A credible test starts the other way around, by translating the utility's objectives into measurable acceptance criteria that are written down before the first sample is collected.
Those criteria have to be specific enough to fail. "Improved solids removal" is not a criterion; a target removal rate, sustained across a defined range of loading conditions, measured by a stated method, is. The same goes for the things that are easy to leave out of a brochure comparison: residual stream characteristics, chemical consumption, reject quality, and the labor the process demands. The discipline of fixing these in advance is exactly why regulators ask for a quality assurance project plan before testing begins. As the EPA puts it in its guidance for treatment studies, a quality assurance project plan is a written document that provides a blueprint for the entire project to ensure it produces reliable data that meet the project's objectives. The plan is not bureaucratic overhead. It is the artifact that keeps the goalposts from moving.
Test under real conditions
Wastewater is not a steady-state feed, and a technology that only ever sees a steady state has not really been tested. Influent strength and flow swing through the day on a predictable diurnal rhythm, with morning and evening peaks several times the overnight low. They swing again, and far more violently, with the weather. The EPA notes that many municipalities experience high influent flows during wet weather, referred to as peak flows, that exceed the treatment capacity of existing treatment units. A process that performs beautifully at average dry-weather flow and then washes out during the first real storm has not earned a place in a design.
So the conditions a test exposes the technology to matter as much as its duration. A good evaluation deliberately captures the range the full-scale plant will actually face: the diurnal load swing, wet-weather hydraulics, seasonal temperature shifts, and the upsets that are routine rather than exceptional. This is why pilot guidance is written around representative loading rather than convenient loading. Minnesota's pilot-testing protocol, for instance, directs that a study run long enough to ensure the treatment units have experienced the representative range of organic and hydraulic loading rates that could reasonably be experienced at the facility, including seasonal changes. Test on the real influent, through its real variability, and the result starts to mean something.
The point of a field evaluation is not to prove a technology can succeed. It is to find the conditions under which it fails, and then decide whether those conditions matter.
Independence and transparency
Who runs the test and who reports the result is not a procedural footnote; it shapes what the result is worth. A vendor evaluating its own technology is not necessarily wrong, but it is structurally conflicted, and a careful reader has to discount the conclusions accordingly. The value of an independent evaluation is that the entity designing the test, collecting the samples, and writing up the numbers has no stake in a particular answer. That separation is the entire premise behind formal verification programs. The EPA's Environmental Technology Verification effort existed precisely because stakeholders needed an independent, objective, and high-quality source of information for sound decision making about commercial-ready technologies, rather than relying on vendor claims alone.
Transparency is the companion to independence. An evaluation a utility can trust shows its work: the sampling plan, the analytical methods, the periods when the unit was offline, and the data that did not flatter the technology along with the data that did. Selective reporting of the good weeks is the quiet way a real result turns into a hopeful one. The honest version reports the full record, including this article's own caveat where it applies. When a published study is authored by a party with an interest in the technology, that interest should be stated plainly so the reader can weigh it; we hold our own work to the same standard below.
From demonstration to design basis
A successful demonstration and a finished design basis are not the same thing, and conflating them is how good pilots still lead to bad outcomes. The job of the evaluation is to produce the specific information a full-scale design actually consumes, which is more than a headline performance number.
Three things in particular tend to separate a design-grade evaluation from a promising one. The first is sustained performance: not a peak result but the level the process reliably holds across the full range of conditions, because the plant has to be sized for the bad days, not the good ones. The second is operational reality: the maintenance, cleaning cycles, chemical dosing, and operator attention the process demands, because a technology that works only when an expert is standing next to it will not work at three in the morning. The third is failure modes: knowing how the process degrades and what it does when it is pushed past its limits, since the design has to include the redundancy and controls that keep an upset from becoming a permit violation. A demonstration that characterizes all three gives engineers something to build on. One that reports only the average effluent quality leaves the hardest design questions unanswered.
A practical checklist
A useful evaluation, stripped to its essentials, covers the following. None of these are exotic; the discipline is in doing all of them rather than the convenient subset.
- Objectives and acceptance criteria fixed in writing before testing, specific enough that the technology can measurably pass or fail.
- Baseline characterization of the plant's own influent and the incumbent process, so the comparison is against reality rather than a textbook.
- Duration long enough to span the representative range of loading, seasonal variation, and at least one real wet-weather event.
- Sampling and QA governed by a written plan: defined methods, sampling frequency, calibration, and chain of custody.
- Scale sufficient to extrapolate honestly to full size, with the scale-up assumptions stated rather than assumed.
- Operations and maintenance observation recorded as data in its own right: chemical use, cleaning frequency, downtime, and operator effort.
- Exit criteria agreed in advance, so the decision to advance, modify, or walk away follows the evidence instead of sunk cost.
The checklist is deliberately mundane. Most evaluation failures are not failures of sophistication; they are a skipped baseline, a test that ran too short, or a goal quietly rewritten to match the data.
Where CWT fits
Independent testing, demonstration, and evaluation of emerging treatment technologies under real-world conditions is one of Caliskaner Water Technologies' four service lines, and the framework above is how that work is done: objectives first, real influent, full record, design-grade output. CWT's own team applied exactly this approach in the first peer-reviewed performance evaluation of a full-scale primary filtration installation, co-authored with George Tchobanoglous of UC Davis and measuring solids and organics capture under real operating conditions over an ongoing demonstration. That is CWT's own work, offered here as a worked example of the method rather than as independent endorsement. For the related challenge of carrying a proven pilot across the gap to full scale, see our piece on the pilot-to-full-scale valley of death. If your utility is weighing an emerging process and wants an evaluation it can defend to a board and a regulator, we are happy to talk.