Marta Kowalska
Marta learned by watching her mother and grandmother rather than by measuring everything. She adjusts batter by sight and texture, and cares first about whether people want another slice.
Learn Chapters 5–8 of Noise by examining the same Judge × Cake × Occasion model through several purpose-built views, then manipulate its sources of variability, add group influence, test your diagnosis, and reset the entire experience whenever you want.
You have been invited to audit a chocolate-cake competition before the final rankings are released. Everyone used the same rubric, yet the scores do not behave as neatly as the organizers expected. Your task is not to pick a winner. It is to determine why competent judges disagree, which disagreement the system should tolerate, and which procedures are manufacturing avoidable variability.
She does, however, reserve the right to disapprove of causal claims made from one mean, one panel, or one suspiciously persuasive senior judge. Gremlins 👹 appear only when a mistake deserves a little mockery; they never reduce the score.
This is a teaching simulation, not a dashboard. The story gives the variables meaning; the laboratory lets you change them.
You will watch the same judging system from four angles. Chapter 5 asks whether the panel is wrong on average or merely inconsistent. Chapter 6 asks where the inconsistency comes from. Chapter 7 holds the judge and cake constant and changes the occasion. Chapter 8 lets judges influence one another and asks whether consensus has improved the judgment or only synchronized it.
Every score belongs to one Judge × Cake × Occasion cell. The Explore laboratory lets you amplify stable severity differences, case-specific interactions, occasion effects, and—only in discussion-first mode—social influence.
Baking keeps the problem emotionally neutral while preserving exactly what Noise needs: shared criteria, professional discretion, repeatable cases, and social influence.
The numerical model is illustrative. It asks when a competition would want competent judges to agree more closely—not whether all human difference should be eliminated.
Mastery is cumulative Test accuracy. Coverage is bounded at 100% across eight concepts. Practice points reward repetition and can keep growing. Learning cycles count complete feedback loops: a solved Scenario run or a finished randomized Test. Mastery remains cumulative Test accuracy. All are stored only in this browser when local storage is available. Nothing is transmitted.
Marta learned by watching her mother and grandmother rather than by measuring everything. She adjusts batter by sight and texture, and cares first about whether people want another slice.
Leo came to baking through formal training. He weighs precisely, tracks temperatures, and treats reproducibility as evidence that a result is deserved rather than lucky.
Nina has baked for decades and prefers darker, less sweet cakes with a denser crumb. She does not regard contemporary sweetness or airiness as neutral standards.
Owen experiments the way some people annotate books: relentlessly. His chocolate cake uses espresso, olive oil, sea salt, restrained frosting, and a deliberately soft crumb.
Sara trusts tested recipes and executes them carefully. Her cake is balanced, clean, and difficult to fault, though few judges find it unforgettable.
Helen has spent years teaching students to diagnose crumb, structure, emulsification, symmetry, and finishing technique. She sees execution errors quickly.
Marcus begins with the eating experience: aroma, flavor development, bitterness, sweetness, and finish. A technically imperfect cake can still win him over.
Priya is interested in whether the result looks controlled and reproducible. Avoidable irregularity matters because she treats it as evidence about the process.
Daniel works with unconventional flavor combinations and is willing to tolerate departures from convention when the departure produces something distinctive.
Claire has judged local competitions for years and uses the upper end of the scale relatively freely. She cares strongly about whether a cake is pleasurable to eat.
Thomas reserves the top of the scale for unusually complete work. He can agree with Claire about ranking while placing the entire field lower.
Separate systematic displacement from unwanted spread.
Calibration score: 7.0. Independent scores: Helen 6.2, Marcus 7.8, Priya 6.4, Daniel 8.1, Claire 7.7, Thomas 5.9.
The group mean is almost exactly right, yet Owen’s outcome depends heavily on which judge he receives.
If the mean is almost perfectly calibrated, would you call this judging system reliable for an individual baker?
Stable rater severity is not the same thing as judge × cake interaction.
Claire and Thomas rank cakes similarly, but Claire uses a consistently higher portion of the scoring scale.
If two judges’ lines are roughly parallel but vertically separated, which part of the disagreement belongs to the judge rather than the cake?
Helen and Daniel can have similar average severity but cross repeatedly when different cakes activate different evaluative priorities.
If their average scores are similar, what does repeated line-crossing tell you that a mean comparison cannot?
Now the same measuring instrument changes over time.
Early tasting: 7.8. After several unusually sweet cakes: 8.5. Later, after technically exceptional entries: 7.3.
Judge = Marcus and Cake = Owen. Only Occasion changes. What source of variability remains available to explain the movement?
Consensus can increase while informational independence decreases.
One panel hears Thomas’s technical criticism first; another hears Daniel’s originality praise first. Both groups become internally consistent, but their final conclusions diverge.
If within-panel spread shrinks in both groups, what additional comparison is needed before concluding that noise fell?
The framework should grow as one nested system.
Systematic displacement from a calibration or criterion.
Stable severity or leniency differences between judges.
Judge × cake interaction: similar means, different case reactions.
Same judge, same case, different occasion.
The model still contains Judge × Cake × Occasion. Instead of making you decode all three spatially, choose the view that best answers the question you are asking.
What best explains the pattern you loaded?
The laboratory waits until you pause, then translates the visual change back into the judging story.
Each card now loads a real laboratory configuration, asks you to diagnose it, and contributes to cumulative practice.
Questions are drawn from a larger bank and scored cumulatively. Wrong answers trigger targeted explanations.
Your mastery profile is stored locally in this browser until you reset it.
The tool is designed for cycling, not one-pass completion. Review the weak distinction, manipulate it once in Explore, then take a fresh randomized test. When you want a completely clean run, use Reset Everything.
This takes about two minutes to orient. Nothing is scored yet.
You are Lexi, the outside reader of the scorecards. The competition organizers do not need you to choose the best cake; they need you to explain why the same judging system produces different answers. Chapters 5–8 become four stages of that audit: measure the error, locate its source, test whether the same judge is stable, then determine what discussion does to independence. The score tracks retrieval; exploration remains consequence-free.
Learned by watching family; values moisture and flavor over perfect geometry.
Measures precisely and prizes reproducibility, structure, and controlled execution.
Prefers darker, less sweet, denser cakes and rejects the idea that current fashion is neutral.
Uses espresso, olive oil, sea salt, restrained frosting, and a soft crumb.
Executes a proven recipe cleanly and consistently, with few obvious defects.
Pastry instructor; notices crumb, structure, symmetry, and execution first.
Food writer; begins with aroma, flavor, balance, and finish.
Competition judge; asks whether the result is reproducible rather than lucky.
Experimental pastry chef; tolerates departures from convention when they work.
Experienced community judge; uses the upper part of the scale relatively freely.
Senior judge; reserves very high scores for unusually complete work.
Every score still belongs to Judge × Cake × Occasion, but you never have to decode those dimensions as perspective. Matrix fixes the occasion; Score Strip fixes one cake; Score Lines compares judge profiles; Occasion Timeline holds judge and cake constant; Panel Comparison adds the Chapter 8 social overlay.
Explore freely, but use the discipline of a small experiment: one change, one prediction, one observed consequence. The story feed will explain the change after you pause, and every event can be dismissed when it has done its job.
Begin in Learn. When you know the cast, the tool will guide you into Explore. After enough distinct manipulations, it will suggest Test. Excellent work, doamnă—now make the judges misbehave scientifically.
This clears mastery, practice points, learning cycles, scenario completions, test history, reading position, onboarding state, narrative cards, and laboratory settings. The page then returns to the opening dedication, followed by the first-run orientation.
This is for debugging and recovery. It does not affect practice points, coverage, mastery, scenarios, or learning cycles.
No diagnostics recorded.