A Lab’s Shift to Staggered Survey Timing Quietly Reshaped Its Diurnal Mood Findings

Aug 10, 2026 By Jonas Eriksen

In the spring of 2022, a mid-sized psychology lab made a seemingly minor change to its daily mood study. Instead of sending participants a single survey at 9 AM, the lab switched to three random windows spread across morning, noon, and evening. The goal was to capture a more representative slice of daily experience. But the change had an unintended consequence: the lab's published diurnal mood curve, which had shown a steady rise through the morning and a dip in the afternoon, flattened and shifted. The peak moved later, the trough disappeared, and the entire pattern looked different. The data hadn't changed. The clock had.

The Same Mood Data, Read Two Ways

Mood is not a fixed trait. It fluctuates through the day, influenced by sleep, caffeine, workload, and social interaction. A 9 AM survey captures one snapshot, an 8 PM survey another. When a lab administers its survey at a single time, it gets a consistent but narrow view. When it staggers times, it averages across the day, but the average is not the same as the curve.

In the lab's earlier protocol, every participant answered at roughly the same hour, usually soon after waking. That timing captured the morning peak in alertness and positive affect. The staggered protocol, by contrast, pulled in responses from people who were mid-commute, post-lunch, or winding down. Each of those moments carries its own mood signature.

The result was a composite that looked like a gentler, flatter version of the original curve. The morning peak softened, the afternoon slump blurred, and the evening recovery nearly vanished. A reader looking at the two figures side by side would conclude they came from different populations. They came from the same people, just at different times.

Researchers who noticed the shift initially dismissed it as noise. But the pattern persisted across multiple cohorts. The lab's own internal replication, run with both protocols in the same week, showed the effect was systematic. The timing decision had become the hidden variable.

When the Clock Becomes a Confound

Time-of-day effects are a well-documented nuisance in affect research. Circadian rhythms influence cortisol, melatonin, and subjective energy. A survey at 7 AM will catch the groggy, pre-coffee state; one at 10 PM may catch the tired but relaxed end-of-day mood. These fluctuations are real, but they are not the target of most studies, which aim for an average or a trait-like measure.

The standard fix is to control for time of day statistically, or to hold it constant by surveying everyone at the same time. The first approach is common in large longitudinal studies; the second is typical in lab-based experiments. But both assume the researcher knows when the relevant mood state occurs. If the study is about daily rhythms, the timing is the variable, not the nuisance.

In this lab, the shift to staggered timing was adopted to reduce the bias of a single convenience time. The idea was that random windows would wash out the idiosyncratic effects of any one hour. Instead, the windows interacted with response rates. People were more likely to answer in the morning when they were fresh, and less likely in the evening when they were busy or tired. The sample composition changed by hour.

That interaction turned the staggered design into a confound. The aggregate mood curve now reflected not just the daily rhythm, but also who happened to respond at each time. The morning respondents were more alert and more positive; the evening respondents were fewer and more fatigued. The apparent flattening was partly an artifact of selective participation.

A Worked Example: The Shift at One Lab

Consider the concrete sequence. The lab originally sent a single daily survey at 9 AM, with a reminder at 11 AM for non-responders. The response rate hovered around 80 percent, and the mood scores showed a clear morning peak. Then the protocol changed.

Participants were randomly assigned to one of three windows: 8–10 AM, 12–2 PM, or 6–8 PM. The intent was to spread the measurement load and reduce the burden of a fixed early alarm. The first week looked fine: scores in each window varied, but the overall mean was similar to the old 9 AM average.

But by the second week, a pattern emerged. Response rates in the morning window stayed high, around 85 percent. The noon window dropped to 70 percent, and the evening window fell to 55 percent. People skipped the evening survey more often, or answered it late, after dinner, when their mood had shifted.

The lab's statistician re-analyzed the data with time of day as a covariate. The original morning peak reappeared once the response bias was accounted for. Without that adjustment, the published curve showed a muted, delayed rise that peaked around 2 PM and declined slowly. The lab's director later admitted the first version of the figure "looked wrong" but was reluctant to change the protocol mid-study.

Beyond the Lab: Staggered Timing in the Wild

The lab's experience is not an isolated anecdote. Similar timing effects have surfaced in other domains, often with real-world consequences. Consider public health surveys that ask about daily habits. If a survey is administered only in the evening, it may overrepresent people who are home and relaxed, missing the stressed commuter or the shift worker. Conversely, a morning-only survey might capture the pre-work rush, inflating reports of anxiety and low energy.

In workplace well-being studies, the timing of a pulse survey can dramatically alter the results. A survey sent at 10 AM on a Tuesday might catch people in a mid-morning lull, while one sent at 3 PM on a Friday could capture end-of-week fatigue. Companies that track employee engagement over time often change their survey schedule without realizing it, inadvertently creating trends that reflect the clock rather than the culture.

Experience sampling methods, which ping participants at random moments throughout the day, are often hailed as the gold standard for capturing mood in situ. But even these designs face timing challenges. The random pings are not truly random if they are constrained to waking hours, and they still depend on the participant's willingness to respond at that instant. If people are more likely to respond when they are bored or idle, the sample skews toward those states, missing the busy, engaged moments that might matter just as much.

The trade-off is between ecological validity and statistical control. A fixed time offers control but risks missing the daily rhythm. A staggered design captures more of the day but introduces response bias. Experience sampling offers the richest data but at the cost of higher participant burden and more missingness. There is no universally correct choice; the best design depends on the research question and the population.

Quantifying the Damage: Effect Sizes and Power

How much does timing actually matter? In the lab's case, the shift in the peak was roughly two hours, moving from around 9 AM to around 11 AM in the unadjusted analysis. The difference in mean positive affect between morning and evening respondents was on the order of half a point on a five-point scale — a modest but not trivial effect. In a typical mood study with a sample of a few hundred participants, such a shift could easily turn a significant morning effect into a null result, or vice versa.

Statistical power is another casualty. When response rates vary by time window, the effective sample size for each time point shrinks. In the lab's evening window, the 55 percent response rate meant that only about half of the assigned participants contributed data. This uneven missingness reduces the precision of the evening estimates, widening confidence intervals and making it harder to detect true differences between morning and evening mood.

Consider a hypothetical replication: a study with 300 participants, evenly split across three time windows. If the evening window has a 50 percent response rate, the effective sample for that window is only 50 people, compared to 85 in the morning. The standard error for the evening mean is roughly 40 percent larger, making it far less likely that a true evening dip would reach significance.

This is not just a theoretical concern. In the lab's own data, the evening scores were so noisy that the apparent flattening was within the margin of error. Only when the statistician adjusted for time of day did the pattern become clear. Without that adjustment, the study would have concluded that mood was relatively stable across the day — a finding that contradicted the lab's earlier work and decades of circadian research.

Alternative Designs and Their Trade-offs

If staggered timing is problematic, what are the alternatives? One option is to keep a fixed time but vary it between participants: assign some to morning, some to afternoon, some to evening, and treat time of day as a between-subjects factor. This avoids the response bias of within-person staggering, but it conflates time with individual differences. People who volunteer for a morning slot may be systematically different from those who prefer evenings.

Another approach is to use a within-person design with multiple fixed times: each participant completes the survey at the same set of times (e.g., 9 AM, 2 PM, 8 PM) on different days. This controls for individual differences and allows a within-person diurnal curve, but it imposes a heavy burden and may lead to habituation or practice effects.

Experience sampling methods (ESM) offer a middle ground. They ping participants at random moments, typically 5–10 times per day, and capture mood in real time. The randomness reduces the bias of fixed times, but the response rate is often lower, and the data are messy. Participants may ignore pings when busy, leading to missingness that correlates with activity level. Some studies compensate by using higher ping frequency, but this increases burden and may alter behavior.

There is also the option of statistical adjustment after the fact. If the sampling times are recorded, analysts can include time of day as a covariate or use multilevel models to estimate the diurnal curve while accounting for response propensity. This is the approach the lab's statistician used, and it recovered the original pattern. But adjustment is only as good as the model, and it cannot fully correct for selection bias if the missingness is non-random.

Each design has its own failure modes. Fixed times risk missing the daily rhythm; staggered times risk response bias; ESM risks missingness and burden. The choice should be driven by the research question. If the goal is to estimate an average daily mood, a fixed time might suffice. If the goal is to map the full curve, staggered or ESM designs are necessary, but they require careful attention to response patterns.

Counter-Arguments: When Staggered Timing Is Fine

It would be unfair to condemn staggered timing outright. In many studies, the bias introduced is small relative to the effect of interest. If the research question is about individual differences in average mood — for example, comparing depressed and non-depressed participants — the time of day may be less critical, as long as the sampling times are balanced across groups. Staggered timing can also reduce the burden of a fixed early alarm, improving compliance and reducing dropout.

Moreover, the response bias may not always be as severe as in the lab's case. In populations with more flexible schedules, such as students or retired adults, evening response rates might be higher. In studies that use automated reminders or incentives, the missingness pattern could be more uniform. The key is to measure and report response rates by time window, so readers can judge the risk.

Some researchers argue that the aggregate curve is actually more ecologically valid than a fixed-time curve, because it reflects the average experience of people across the day, weighted by their availability. A person who is too busy to answer in the evening is, by definition, not experiencing that moment in a relaxed state; their absence from the evening sample may be informative. This perspective treats the missingness as part of the phenomenon rather than a nuisance.

However, this argument only holds if the missingness is informative about the construct of interest. If the goal is to estimate the diurnal rhythm of mood in the general population, then excluding the busy evening person biases the estimate upward for evening mood. If the goal is to estimate the average mood of people who are available to answer, then the bias is less problematic. The researcher must decide which target is more relevant.

What the Data Can and Cannot Say

Staggered sampling is not inherently flawed. It can reduce the bias of a single convenience time, especially when the target is an average daily mood rather than a precise curve. Many large-scale studies, including national well-being surveys, use random or stratified sampling times precisely because they want a representative snapshot across the day.

The problem arises when the sampling times are not random with respect to the outcome. If response propensity varies by hour, and mood varies by hour, the two are confounded. The resulting curve is a mix of the true diurnal rhythm and the response pattern. No statistical adjustment can fully separate them without strong assumptions.

Reasonable people disagree about the best solution. Some argue for fixed times with rigorous compliance checks; others prefer random times with high-frequency sampling, such as experience sampling methods that ping participants at random moments. The latter approach captures more of the day but also more missing data, since people cannot always respond instantly.

The lab's experience illustrates a broader point: the sampling window is not a neutral choice. It is a hypothesis about when the effect of interest lives. A study of morning mood should not survey at midnight, and a study of evening wind-down should not rely on 9 AM responses. The design must match the construct.

Why Methodology Stories Matter

Science is a process, not a set of findings. Every published result carries the fingerprints of the procedural choices that produced it, from the sample size to the survey instrument to the time of day. These choices are rarely visible in the final paper, which typically reports the outcome and the statistical test, not the logistics.

The staggered timing shift is a reminder that the most influential variable in a study may be the one the researchers did not intend to manipulate. The lab's original 9 AM protocol had its own bias, but it was a stable bias. The staggered protocol introduced a variable bias that shifted with the clock.

Peer reviewers rarely catch such artifacts. They check for statistical power, multiple comparisons, and p-values, but they seldom ask when the data were collected. The versioned plot archives that some ecology groups now maintain are a step toward transparency, but they do not capture the timing of each measurement.

Readers who want to evaluate a mood study should look past the headline and into the methods section. The details matter: the exact wording of the survey, the response window, the reminder schedule, and the time zone. Each is a choice that could flip the result.

Practical Checks for Any Mood Study

First, ask when the surveys were administered. A single fixed time is easier to interpret but may miss the daily rhythm; staggered times capture more but introduce response bias. Look for a table that reports response rates by time window.

Second, check whether the timing varied across participants. If some people answered in the morning and others in the evening, the study is comparing different hours as well as different people. The authors should have controlled for time of day in the analysis.

Third, re-analyze the data with time as a covariate, as the lab's statistician did. If the mood curve changes substantially, the timing is a confound. The original paper, without that adjustment, would have told a different story.

Finally, demand transparency. A good methods section reports the sampling windows, the response rates per window, and the rationale for the timing plan. The voxel size in a neuroimaging study can shift activation maps; the survey clock can shift mood curves just as decisively.

The lab's shift was not a mistake. It was a trade-off, and the trade-off had a cost. The next study should weigh that cost explicitly, before the data are collected, not after the curve has flattened.

Recommend Posts
Science

A Nanoparticle’s Size Distribution, Not Its Chemistry, Drove One Catalyst’s Turnover Gap

By Jonas Eriksen/Aug 10, 2026

Two labs reported conflicting catalyst turnover numbers despite identical chemistry. The gap traced to nanoparticle size distribution, not composition. A methodology explainer.
Science

A Moth Surveyor’s 1970s Light Trap Grid Now Calibrates Urban Bat Detectors

By Renu Shah/Aug 10, 2026

How a 1970s moth survey grid now calibrates urban bat detectors, improving acoustic monitoring reliability through cross-disciplinary method borrowing.
Science

A Salmon Louse’s Genome Draft Sat Uncited for Years Until One Lab Rebuilt Its Reference

By Alice Chen/Aug 10, 2026

A fragmented salmon louse genome draft sat uncited for years. One Norwegian lab's meticulous rebuild turned it into an indispensable reference, reshaping parasite genomics.
Science

The Calcium Signal’s 40-Hertz Tag Confirmed Only After One Lab Switched Its Behavioral Scoring

By Jonas Eriksen/Aug 10, 2026

How a single lab's switch from manual to automated behavioral scoring turned a shaky 40-Hz calcium signal into a robust finding, with lessons for neuroscience.
Science

A Carbon Observatory’s Ancillary Weather Station Outlasted Its Main Spectrometer’s Funding

By Jonas Eriksen/Aug 10, 2026

A carbon observatory's main spectrometer lost funding, but its cheap weather station kept running, proving that low-cost ancillary data can outlast expensive science.
Science

Molybdenum Disulfide’s 2018 Conductivity Claim Faltered When Three Labs Retested Its Crystal Purity

By Karim Osman/Aug 10, 2026

A 2018 claim of near-metallic conductivity in MoS2 crystals failed when three labs retested purity, found contaminants, and couldn't replicate the results.
Science

A Radio Telescope's Sea-Cliff Siting Outlived Two Decades of Its Receiver Upgrades

By Jonas Eriksen/Aug 10, 2026

A sea-cliff radio telescope's location has outlasted two decades of receiver upgrades. The quiet-zone advantage and horizon access prove that siting physics often outweighs hardware improvements.
Science

A Polymer Batch's Drying Oven Setpoint, Not Its Recipe, Determined One Lab's Mechanical Test Spread

By Jonas Eriksen/Aug 10, 2026

A polymer lab's tensile test scatter traced back to the drying oven's setpoint, not the recipe. This methodology feature explores how an overlooked thermal step shaped mechanical outcomes and what it means for reproducible materials science.
Science

A Lab’s Shift to Staggered Survey Timing Quietly Reshaped Its Diurnal Mood Findings

By Jonas Eriksen/Aug 10, 2026

How a lab's shift from fixed to staggered survey timing quietly altered its diurnal mood curve, turning a procedural choice into a hidden variable.
Science

Thirty Years of Duty-Cycle Logs Show One Telescope’s Dome Cost Exceeds Its Detector’s Own Budget

By Karim Osman/Aug 10, 2026

A look at how three decades of duty-cycle logs reveal that dome operations can outpace detector budgets, and why observatory funding rarely accounts for this.
Science

Sea-Surface Temperature Proxies From 2,000 Foraminifera Shells Pinpoint the 1910s Warming Onset

By Alice Chen/Aug 10, 2026

A study of 2,000 foraminifera shells uses magnesium-to-calcium ratios to trace sea-surface temperatures, pinpointing the 1910s as a key warming onset. The method and its limits explained.
Science

A Two-Pound Beaker Weighing Protocol Split One Lab’s Oxygenesis Replication

By Jonas Eriksen/Aug 10, 2026

A contested microbial metabolism claim split labs. The culprit: a two-pound beaker and a weighing protocol that varied. Here's how mundane details derailed replication.
Science

A Field Team’s Decision to Tag 400 More Deer Overturned a Predator-Prey Model

By Karim Osman/Aug 10, 2026

A field team's decision to tag 400 more deer on Isle Royale overturned a long-standing predator-prey model, revealing a Type III functional response and reshaping wildlife management.
Science

A Palladium Membrane’s Hydrogen Permeability Data Recalibrated Fuel Cell Anode Models

By Alice Chen/Aug 10, 2026

New measurements of palladium membrane hydrogen permeability challenge decades-old constants, reshaping fuel cell anode models and cost estimates.
Science

The Replication Crisis’s Career-Spanning Data Finally Reached Economists’ Field Experiments

By Jonas Eriksen/Aug 10, 2026

How the replication crisis that shook psychology finally reached economics' field experiments, what it revealed about effect sizes, and how pre-registration and open data are changing the field.
Science

Neuropixels Probe Rental Fees Now Eclipse One Lab's Animal Housing Budget

By Jonas Eriksen/Aug 10, 2026

Rental fees for Neuropixels probes now rival or exceed animal housing costs in some labs, reshaping budgets and research planning.
Science

A Data Descriptor's Mandatory Code Deposit Unearthed a 2011 Climate Model's Hidden Calibration Choice

By Alice Chen/Aug 10, 2026

A mandatory code deposit in a data descriptor revealed a hidden calibration choice in a 2011 climate model, exposing gaps in reproducibility and uncertainty estimates.
Science

A Two-Photon Laser’s Beam Waist Recalibration Reversed One Lab’s Dendritic Spine Counts

By Renu Shah/Aug 10, 2026

A routine beam waist recalibration reversed a lab's dendritic spine counts, revealing an optical artifact mistaken for biological change. A lesson in optical hygiene.
Science

Darwin’s Beak Measurements, Replotted by Hand, Flipped One Grant’s Speciation Verdict

By Alice Chen/Aug 10, 2026

A graduate student's hand-plotting of the Grants' finch data uncovered a bimodal beak distribution, prompting a reanalysis that refines, not overturns, the original speciation interpretation.
Science

The 1976 Code-Archiving Mandate That Outlived Its Telescope’s Entire Optics Budget

By Alice Chen/Aug 10, 2026

How a 1976 code-archiving rule from a federal funder outlasted its telescope's optics budget, shifting costs to researchers and shaping today's reproducibility push.