// DOE Masterclass

Design of Experiments,
for scientists who run experiments

Design of Experiments for bench scientists - taught the way you plan real experiments, not the way statistics textbooks teach it.

8 video modules~8 minutes totalNo statistics prerequisiteFree
// Why this course exists

Most experiments fail before the bench

01 ·

OFAT is not rigor

Varying one factor at a time feels careful - and is structurally blind to interactions. Multifactorial designs learn more from fewer runs. That is the whole course in one sentence.

02 ·

No statistics prerequisite

Every concept is taught the way you plan real experiments - factors, plates, batches, noise - not through distribution theory. If a formula appears, it earns its place.

03 ·

Applied, not academic

Module 07 takes a real research question through Shadow AI - literature, ranked hypotheses, and a bench-ready DOE with controls, materials, and a statistical plan.

// Curriculum

8 modules, narrated by a scientist

Short on purpose - each module is a few minutes of signal, watched between incubations. The order matters: screen before you optimize, power before you pipette.

01 ·WHY DOE0:52

One factor at a time is the slowest way to learn

Most labs still vary one factor at a time - OFAT. It feels careful, but it is statistically the weakest and most expensive way to explore a system. This module shows why multifactorial designs find interactions OFAT is structurally blind to, with fewer runs.

  • Why OFAT misses every interaction by construction
  • What a factorial design actually measures
  • The run-count arithmetic - when 12 designed runs beat 30 sequential ones
Read the transcript & key terms

For over a decade at the bench, I planned experiments the way most of us were taught - change one factor, hold everything else still. It feels careful. It is actually the slowest way to learn.

One factor at a time has a structural blind spot - interactions. If pH only matters at low temperature, OFAT will never see it, because it never varies the two together. Real biological systems are full of these couplings.

A factorial design varies factors together, on purpose, in a balanced pattern. Every run contributes information about every factor - so twelve designed runs routinely tell you more than thirty sequential ones. Not because of statistical tricks - because of geometry.

That is the whole course in one sentence. Stop asking one question per experiment. Design the experiment to answer several - at once, with the same pipettes, in fewer runs. In the next module, we frame your experiment as a design space.

OFAT
Changing one factor at a time while holding the rest fixed. Structurally cannot detect interactions, because it never varies two factors together.
Factorial design
A design that varies factors together in a balanced pattern, so every run carries information about every factor.
Interaction
When the effect of one factor depends on the level of another - e.g. pH mattering only at low temperature.
02 ·THE DESIGN SPACE0:53

Factors, levels, responses - your experiment as a design space

Before any design can be chosen, the question has to be framed: what can you turn (factors), how far can you turn it (levels and ranges), and what will you measure (responses). Getting this frame right is most of the work - and most of what goes wrong.

  • How to choose factors and defensible ranges
  • Continuous vs categorical factors - and why it changes the design
  • Picking responses you can actually measure with acceptable noise
Read the transcript & key terms

Before choosing any design, frame the question. Three decisions do most of the work.

First - factors. The variables you can actually turn. Temperature, pH, concentration, media, strain. Be honest about what is controllable at your bench, not in principle.

Second - levels and ranges. How far you will turn each one. Too narrow, and effects hide inside the noise. Too wide, and you leave the region where the chemistry even works. A defensible range comes from the literature, and from what you already know.

Third - responses. What you will measure. A good response is quantitative, relevant to the goal, and measurable with noise you can live with. Yield, titer, activity, viability.

Together these define a design space - a geometric region your experiment will explore. Every design in this course is just a pattern of points placed in that space. Frame the space well, and the rest is almost mechanical.

Factor
A variable you deliberately change: temperature, pH, concentration, media, strain.
Level
A specific setting of a factor. Two-level designs use a low and a high.
Response
What you measure - yield, titer, activity, viability.
Design space
The geometric region defined by your factors and their ranges; every design is a pattern of points inside it.
03 ·SCREENING0:55

Screening designs - find the vital few

With 6-10 candidate factors you do not optimize - you screen. Plackett-Burman and fractional factorial designs rank main effects in remarkably few runs, so the expensive follow-up work is spent only on factors that matter.

  • When Plackett-Burman is the right tool - and its aliasing cost
  • Fractional factorials and design resolution (III, IV, V)
  • Reading a screening result: effect ranking, not final answers
Read the transcript & key terms

With eight candidate factors, you do not optimize - you screen. The goal is ranking, not precision. Which factors actually move the response?

Plackett-Burman designs are the extreme case - twelve runs can screen up to eleven factors. The price is aliasing. Main effects are entangled with interactions. That is acceptable at this stage, because all you want is the shortlist.

Fractional factorial designs are the finer instrument. You run a deliberate fraction of the full factorial - a half, or a quarter - and the design's resolution tells you exactly what is entangled with what. Resolution four keeps main effects clean of two-factor interactions.

Read a screening result as an effect ranking - a Pareto chart of what matters. Two or three factors usually dominate. Those go forward to optimization. The rest get fixed at sensible values, and stop costing you runs. Screening is how DOE pays for itself first.

Screening design
A design that ranks which factors matter, rather than estimating effects precisely.
Plackett-Burman
A screening design that handles up to k-1 factors in k runs (k a multiple of four). Cheap, but main effects are aliased with interactions.
Fractional factorial
A deliberate fraction - a half, a quarter - of the full factorial.
Aliasing
When two effects cannot be separated by the design and appear as one combined estimate.
Resolution
How badly a fractional design aliases. Resolution IV keeps main effects clear of two-factor interactions.
04 ·OPTIMIZATION0:58

Response surface methods - find where it is best

Once screening has cut the field to 2-4 factors, response surface designs - central composite, Box-Behnken - map curvature and locate optima. This is where DOE stops saving runs and starts finding conditions you would not have tried.

  • Why optimization needs center points and axial runs
  • Central composite vs Box-Behnken - choosing between them
  • Reading a response surface: optima, ridges, and trade-offs
Read the transcript & key terms

After screening, you hold the two or three factors that matter. Now the question changes - from which, to where. Where in the design space is the response best?

Straight-line models cannot answer this. Optima live in curvature, and seeing curvature takes more than two levels. Response surface designs add exactly the runs that reveal it. Center points measure it. Axial points map it.

The central composite design is the workhorse - a factorial core, axial runs, and repeated center points. Box-Behnken is the alternative, when the corners of your space are dangerous or expensive to reach.

What comes back is a response surface - a fitted map of your system. Read it for three things. The optimum. The ridges, where trade-offs live. And the plateaus, where the process is robust. Often the surface points somewhere just outside your ranges. That is not failure - that is the map telling you where the next design goes.

Response surface
A fitted model of how the response varies across the design space, capable of showing curvature.
Central composite design
A factorial core plus axial (star) runs plus repeated centre points - the workhorse optimisation design.
Box-Behnken
An alternative response-surface design that avoids the extreme corners of the space.
Centre point
A run at the midpoint of every factor. Measures curvature and gives an estimate of pure error.
05 ·RIGOR0:54

Replication, randomization, blocking - the three defenses against noise

Biology is noisy; the design has to defend itself. Replication estimates noise, randomization protects against drift and confounding, blocking removes the variation you already know about - day, plate, batch, operator.

  • Technical vs biological replicates - what each one buys you
  • What randomization actually protects against (it is not superstition)
  • Blocking on plate, day, and batch without burning runs
Read the transcript & key terms

Biology is noisy. A design that ignores that will find effects that are not there - and miss the ones that are. Three defenses, in order.

Replication estimates the noise. Without replicates, you cannot tell an effect from a fluctuation. And know the difference between technical replicates - the same sample, measured again - and biological replicates, which capture the variation you actually care about.

Randomization protects against what you cannot see - drift in the instrument, the afternoon warming of the lab, the order your hands tire in. Run order is randomized not from superstition, but so hidden trends cannot masquerade as effects.

Blocking removes the variation you already know about. Day, plate, batch, operator. Let each block absorb its own baseline, and the factor effects come out cleaner.

Replicate to measure the noise. Randomize against the unknown. Block the known. Every good design leans on all three.

Technical replicate
The same sample measured again. Captures measurement noise only.
Biological replicate
An independently prepared sample. Captures the variation you actually care about.
Randomisation
Running conditions in random order so unseen drift cannot masquerade as a factor effect.
Blocking
Grouping runs by a known nuisance source - day, plate, batch, operator - so its variation is removed from the comparison.
06 ·STATISTICAL POWER0:50

Power and sample size - how many runs you actually need

Underpowered experiments are the quiet failure mode of bench science - they cannot see the effect they were built to find, and the null result gets believed anyway. This module makes the effect size / noise / run count trade-off concrete.

  • Effect size, variance, and power - the triangle you cannot cheat
  • Rules of thumb for factorial designs that hold up in practice
  • What to do when the honest answer is 'more runs than you can afford'
Read the transcript & key terms

The quiet failure mode of bench science is the underpowered experiment - a design too small to see the effect it was built to find. The result reads as no effect. It gets believed. It is wrong.

Power lives in a triangle you cannot cheat: the effect size you care about, the noise of your response, and the number of runs. Fix any two, and the third is decided for you.

The honest sequence goes like this. Estimate your noise - from pilot data, or history. Name the smallest effect worth acting on - a ten percent gain in titer, one log in viability. Then compute the runs. For factorial designs, a useful rule of thumb - every effect you want to see clearly wants the equivalent of eight to sixteen runs behind it.

And when the honest answer is more runs than you can afford - shrink the question, not the rigor. Fewer factors, screened well, beat many factors measured badly.

Statistical power
The probability of detecting an effect that is genuinely there. 80% is conventional.
Effect size
The difference worth acting on, often standardised by the response's standard deviation.
Underpowered
A design too small to detect the effect it was built to find; its null result carries little information.
07 ·THE WALKTHROUGH1:16

From research question to DOE in Shadow AI

The product module. A real research question goes in; Shadow AI runs the literature across PubMed, OpenAlex, Semantic Scholar and arXiv, proposes ranked hypotheses, and returns a bench-ready design - factors, levels, controls, materials, statistical plan - in minutes. Everything from modules 01-06, applied.

  • Describing a research problem so the agent designs the right experiment
  • Reviewing generated hypotheses and the literature behind them
  • Reading the design output: DOE table, controls, materials, statistical plan
Read the transcript & key terms

Everything so far - applied. Let me show you what this looks like in Shadow AI.

I start with a plain-language research question. The same sentence I would say to a colleague. Which buffer conditions maximize the stability of my protein formulation? That is all the framing the agent needs.

Shadow AI reads the literature first - PubMed, OpenAlex, Semantic Scholar, and arXiv - and comes back with what is already known. Which factors have moved this response before, in whose hands, at what ranges.

Then it proposes hypotheses - ranked, falsifiable, each traceable to its citations. I review them the way I would review a student's. Accept, edit, discard.

Then, the design. Factors and defensible ranges. A screening or response-surface design chosen for the question - fractional factorial, Plackett-Burman, or central composite - with the reasoning stated. Controls - positive, negative, vehicle. Replication and blocking, laid out on the plate. A materials list, with quantities and calculations. And a statistical plan written before the first sample is run - which is exactly the discipline this course has been teaching.

What used to take me ten hours of planning is now a few minutes of structured review. The thinking is still mine. The agent just does the assembly. The Explorer plan is free - bring your own question, and watch it become a design.

Bench-ready design
A design complete enough to run: factors and ranges, the DOE, controls, replication and blocking, materials, and a statistical plan.
Falsifiable hypothesis
A hypothesis stated so that a specific result would disprove it.
Statistical plan
How the data will be analysed, written before the first sample is run.
Follow along in Shadow AI - free
08 ·ANALYSIS & ITERATION0:57

Reading the results - and designing the next campaign

A DOE is rarely one experiment - it is a campaign. This module covers what to look at first in the analysis (effects, then model fit, then residuals), what a 'failed' design still teaches you, and how the result of one design seeds the next.

  • Effects and interactions first - p-values last
  • Model diagnostics a non-statistician can and should check
  • Sequential experimentation: screen, optimize, confirm
Read the transcript & key terms

The design was the hard part. The analysis, read in the right order, is almost calm.

Effects first. Which factors moved the response, by how much, in which direction - and which interactions mattered. This is the answer to the question you designed. P-values come last, not first. They qualify the answer - they are not the answer.

Then, model fit. R-squared in context, and the residuals. Do they look like noise - or like a pattern the model missed? A curved trend in the residuals means curvature you have not modeled yet. That is information, not failure.

Then, the decision. A DOE is rarely one experiment - it is a campaign. Screen, optimize, confirm. A screening result seeds the response surface. A surface optimum gets a confirmation run. Even a design that finds nothing has told you where the effect is not - and spared you a year of chasing it.

Design. Run. Read. Design again. That is the method. The next experiment you plan is module nine.

Main effect
The average change in response when a factor moves from low to high.
Residual
The difference between an observed value and the model's prediction. Patterns in residuals mean the model is missing something.
Sequential experimentation
Screen, then optimise, then confirm - a campaign of designs rather than one experiment.
// Keep going

Put a number on your own experiment

Module 06 covers power. The free calculator does the arithmetic for your design - replicates per group, the smallest effect a fixed run budget can detect, and the run count for a factorial, Plackett-Burman or central composite design.

Open the sample size calculator →
// Your instructor

Taught by someone who lived the planning chaos

Ankita Pandey
Ankita Pandey
Founder and CEO, Shadow AI

Over a decade as a bench scientist in pharmaceutical R&D, including Novartis and Momenta. Ankita built Shadow AI after living the planning chaos this course teaches you to escape - she narrates every module.

LinkedIn ↗

Watch the course, then design one for real

The Explorer plan is free - 3 complete experiment design cycles a month. Bring the research question you have been putting off.