Skip to content
A second look at the obvious answer
Think Twice TodayA second look at the obvious answer

Habits

Running an experiment on yourself without fooling yourself

A personal trial can genuinely settle whether something works for you, and the same features that make it feasible are the ones that make a confident wrong answer so easy to reach.

By Samar Bhatia4 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Why the sample of one isn’t automatically worthless

Averages from a large trial answer a question about a population, and a person isn’t a population. Where responses to something genuinely vary between individuals, the average can be an accurate description of the group and a poor guide to you, which is the strongest argument for testing something on yourself rather than reading about it.

The catch is that a self-experiment concentrates every source of bias into a single observer who is also the subject, the analyst, and the person with an interest in the outcome. Nothing about being your own participant removes the problems a study design exists to solve; it removes the personnel who would otherwise be handling them.

The four things that will produce a false positive

Expectation is the first. Believing a change will help produces reported improvement in almost anything measured subjectively, which isn’t imagination but a real and reliably observed effect on what people report and sometimes on what they do.

Regression is the second and it’s the quiet one. People start interventions when things are unusually bad, and unusually bad measurements are followed by less extreme ones for purely statistical reasons. Whatever you started on a terrible week will look effective the following week.

The third is that you changed several things at once, which almost everybody does, because starting something new tends to come with sleeping differently, paying more attention, and a general burst of effort. The fourth is that you decided what counted as success after seeing the results, which is the same analytic flexibility that inflates the published literature and works just as well on a spreadsheet of your own numbers.

A design that handles most of it

Write down, before starting, the single thing you’re changing, the single measure you will use, how often you will record it, and how long the trial runs. Committing to the measure in advance is the highest-value step, because it removes the freedom to notice afterwards which of several outcomes moved.

Then take a baseline for long enough to see the ordinary variation, which is usually longer than people expect. Without knowing how much the measure bounces around on its own, you cannot tell whether a change afterwards is anything at all. Most self-experiments fail here rather than at any sophisticated statistical step.

If the intervention can be started and stopped, alternate blocks of doing it and not doing it, several times, in an order set in advance. Repeated alternation is what distinguishes a real effect from a trend, and it is the closest a single person can get to a control condition.

What cannot be fixed and should be admitted

You cannot blind yourself to most things, which means expectation stays in the results permanently. For anything measured by how you feel, that is a serious limitation and the honest response is to report the finding as what it is: this arrangement is associated with feeling better, by a person who expected it to.

Nor can you generalise the result to anybody else, or often to yourself in a different season or circumstance. A finding from one person over six weeks is a finding about one person over six weeks. That is genuinely useful for the decision you were making and it is not knowledge about the thing being tested.

Anything touching health belongs with a clinician who knows your situation. A personal trial is a reasonable way to decide about a working routine or a daily arrangement, and it is not a substitute for advice about a medical question, where the ways of being wrong are less forgiving.

Reading your own results honestly

Compare the change against the ordinary variation you measured at baseline, not against zero. If the measure normally swings by a certain amount week to week, a movement of that size after an intervention is not a result, and treating it as one is how people end up committed to arrangements that never did anything.

And decide in advance what would count as a failure, in numbers, because deciding afterwards is where most of these exercises quietly collapse. A trial with no stated failure condition cannot produce a negative result, which means it cannot produce a positive one either — it can only produce a confirmation.

The value that survives all the caveats

Even a badly controlled personal trial does something a decision made from reading cannot: it forces the vague question of whether a thing helps into a specific measure over a specific period. A great deal of the benefit comes from that forcing rather than from the result, since it usually turns out that nobody had defined what improvement would look like.

The other benefit is that a written record of what you tried and what happened accumulates. One trial is weak evidence; a folder of them, run the same way over a few years, is the only calibrated information about yourself that anybody is ever going to have, and it does not exist unless somebody writes it down at the time.

Common questions

How long should a personal trial run?

Long enough to see the measure vary on its own before you start, and long enough afterwards that a single unusual week cannot dominate the result. There is no fixed answer, and the baseline period is the part people cut short.

Can I test more than one change at a time?

You can, and you will not learn which one did anything. If you only care whether the package works and never need to drop a component, that may be acceptable. If you want to know what to keep, the changes have to be separated in time.

What if I cannot measure the thing I care about?

Then choose a proxy deliberately, write down what it is standing in for, and note what would make the proxy move without the real thing moving. That note is what lets you notice later that you optimised the measurement rather than the outcome.

Habitsself-experimentmeasurementevidencemethod
Samar Bhatia
Editor, Think Twice Today

Samar joined to cover biases, choices, risk and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.