Rularoo

The experiment behind Rularoo

The 2-4-6 task

In 1960 Peter Wason showed people three numbers and asked them to work out the rule behind them. Most of them never did, and the way they failed is the reason this game exists.

What Wason actually did

Twenty-nine students were shown the sequence 2, 4, 6 and told it obeyed a rule the experimenter had in mind. They could write down any three numbers they liked, as often as they liked, and each time they were told whether that sequence obeyed the rule. When they were certain, they announced the rule itself.

Almost everyone had a theory within a minute: add two each time. Then they tested it. 8, 10, 12. Yes. 14, 16, 18. Yes. 20, 22, 24. Yes. Three confirmations was usually enough to feel certain, and the theory was wrong.

Run the experiment on yourself

This is the 1960 task, unchanged, with the rule Wason used. Test as many sequences as you like. When you are sure, announce the rule.

This sequence obeys the rule

  1. 2
  2. 4
  3. 6

Three whole numbers, separated by commas.

    Announce the rule

    Your rule reads

    Why confirming cannot work

    The theory people reach for, add two each time, only accepts sequences that the real rule also accepts. It sits entirely inside the truth. So every sequence chosen to confirm it comes back YES, including the ones the real rule was accepting for a completely different reason. The agreement is guaranteed, which is exactly what makes it worthless.

    The only test that carries information is one your own theory says should fail. If you believe the rule is add two each time and you try 1, 2, 100, your theory predicts NO. When the answer comes back YES you have learned in one move what a hundred confirmations could never have told you.

    Klayman and Ha argued in 1987 that this is less a bias than a habit. People use a positive test strategy, checking the cases where they expect their hypothesis to hold, which is a sensible default almost everywhere in life and a disaster on this particular task. The behaviour is the same either way.

    What Rularoo does with it

    Wason's task has one rule and you meet it once. Rularoo generates a new one every day at three difficulties, judges your answer by how it behaves rather than how you worded it, and scores you against par: the number of probes a perfect player would need if every test split the remaining possibilities as evenly as it could.

    Par is the part that changes the lesson. In 1960 the question was whether you could find the rule at all. Here it is findable and the question is how efficiently, and par is unforgiving about confirmation, because a probe whose answer you already know splits nothing and still costs you a stroke.

    Play today's puzzle Learn the rules

    Using this in a classroom

    The 2-4-6 task is a standard demonstration in psychology, statistics and scientific method courses, and it works because students fail it in front of themselves instead of being told that they would. This page is free, needs no account and never asks for one. Run the experiment above as a group, then set the daily puzzle as the follow-up: the same lesson, but scored, so the habit has somewhere to go.

    Sources

    Wason, P. C. (1960). On the failure to eliminate hypotheses in a conceptual task. Quarterly Journal of Experimental Psychology, 12(3), 129-140.

    Klayman, J., and Ha, Y.-W. (1987). Confirmation, disconfirmation, and information in hypothesis testing. Psychological Review, 94(2), 211-228.

    Rularoo is better in the app

    Every puzzle here is free to play in the browser, today's and the whole archive. The app is for the part a browser cannot do: a puzzle waiting for you every morning, a streak that survives a missed one, and the Coach telling you which probe cost you the round.