← Back to report

FLOWDEALER · METHODOLOGY

Methodology · The Execution Paradox · Q3 2026

N=603 responses were collected; one was a junk response, so all results are computed over the 602 analysable ones, which were collected between 2026-01-18 and 2026-03-07. Recruitment was targeted rather than open: 517 respondents (86%) came through the Prolific participant panel against a screen described in full below, 33 through launch partners, 14 through an organization and 3 through other channels, with 35 carrying no recorded source. The sample is side-builders, founders, freelancers, business owners and employees (30.6% / 20.3% / 17.6% / 15.1% / 12.1% respectively, plus 4.2% on early startup teams and one respondent who described themselves differently). Data-quality checks flagged 4 of 602 possible straightliners (0.7%, within-respondent SD below 0.5 across 11 numeric items); they were retained, and removing them moves no published mean by more than 0.004. The burnout composite (8 items, each normalized to 0 to 1 and the average rescaled to 0 to 5) showed acceptable internal consistency, Cronbach's α = 0.797. Bootstrap 95% percentile intervals (2,000 resamples) were used for means and Wilson score intervals for proportions. Analysis is complete-case per metric with no imputation. Results describe this specific screened sample, not the general working population.

01 · SAMPLE

Sample

How the sample describes its work situation (602 of 602 respondents)
Work situation%Respondents
Side builder30.6%184
Founder20.3%122
Freelancer17.6%106
Business owner15.1%91
Employee12.1%73
Early startup team4.2%25
Described themselves otherwise0.2%1
Where the 602 respondents came from
ChannelRespondents% of sample
Prolific participant panel51785.9%
Launch partners335.5%
No recorded source355.8%
Through an organization142.3%
Reddit, a Tally form, other30.5%

The panel recruitment was screened, not open. It applies to the 517 respondents recruited through the participant panel, who had to meet every one of these criteria:

  • fluent in English
  • in full-time employment
  • within an age band (23 to 45, and 25 to 43 on some batches)
  • holding a job position, with a stated number of years of work experience
  • using AI at work weekly or more often
  • doing at least a quarter of their work on a computer
  • a stated level of education
  • a stated employer or organization type
  • currently running something entrepreneurial
  • a panel approval rate of 99 to 100
  • not a respondent to any earlier batch, so nobody answered twice

The remaining 85 respondents came through Flowdealer's own channels and were not screened on these criteria.

02 · COHORT CROSS-TABS

Cohort cross-tabs

Say-do-gap flags against burnout (602 respondents, every flag level shown)
Say-do-gap flagsRespondents% of sampleMean burnoutScored
032053.2%1.47316
116427.2%2.12162
26611.0%2.4366
3386.3%2.9637
4101.7%3.1710
530.5%3.493
610.2%3.751
Hours per week against outcomes (602 respondents; the two longest brackets hold 17 and 9 people, so read them as directional)
Hours per weekRespondentsProductive bestCompletionQuality of lifeBurnout
<40942.833.513.001.89
40-603483.113.713.121.79
60-801343.073.512.632.11
80-100173.003.532.762.16
100+93.893.442.782.16
03 · INSTRUMENTS

Instruments

Survey instrument, all 25 questions
ColumnQuestion as askedResponse formatDirection
q1_rolesWhat roles do you perform?Select any of eleven named functions, plus free entryNeither
q2_work_situationWhat's your work situation?Side builder · Founder · Freelancer · Business owner · Employee · Early startup team, plus free entryNeither
q3_productive_bestHow close to your "productive best" are you right now?Slider, 0 to 5Higher is better
q4_quality_of_lifeHow would you rate your current quality of life?Slider, 0 to 5Higher is better
q5_hours_per_weekHow many hours per week do you spend working?Under 40 · 40 to 60 · 60 to 80 · 80 to 100 · 100 or moreNeither
q6_plans_scheduleDo you actively plan your schedule?Yes, regularly · Sometimes, inconsistently · No, I don't planNeither
q7_replan_frequencyHow often do you have to re-plan your day due to unexpected changes?Almost never · Few times a month · Few times a week · Almost daily · DailyHigher is worse
q8_textWhat helps you follow through on your commitments and responsibilities? What doesn't work?Free textNeither
q9_external_pressure_neededHow often do you find it difficult to follow through without external pressure?Almost never · Few times a month · Few times a week · Almost daily · DailyHigher is worse
q10_priority_difficultyHow often do you find it difficult to decide what your next priority should be?Almost never · Few times a month · Few times a week · Almost daily · DailyHigher is worse
q11_overcommit_frequencyHow often do you commit to more than you can actually deliver?Almost never · Few times a month · Few times a week · Almost daily · DailyHigher is worse
q12_plan_completionBy the end of a typical week, how much of what you planned actually gets done?Slider, 0 to 5Higher is better
q13_textWhen executing your plan is harder than it should be, what's usually the reason?Free textNeither
q14_primary_drainWhich drains you most on a typical workday?Cognitive load · Social interactions · Emotional demands · Physical fatigue · All drain me roughly equallyNeither
q15_textIf you could fix one thing about your work ethic or execution, what would it be?Free textNeither
q16_tool_countHow many tools do you actively juggle to manage work and life?1 to 3 · 4 to 6 · 7 to 9 · 10 or moreNeither
q17_tools_usedSelect your 3 to 5 most frequently used tools for managing workSelect from a named tool list, plus free entryNeither
q18_textWhat do you dislike the most about your current tools and systems?Free textNeither
q19_driveHow hungry are you to achieve more?Slider, 0 to 5Higher is better
q20_wellbeing_priorityHow important is it that success doesn't cost you your health or sanity?Slider, 0 to 5Higher is better
q21_overwhelm_frequencyHow often do you feel overwhelmed by life, commitments and/or responsibilities?Almost never · Few times a month · Few times a week · Almost daily · DailyHigher is worse
q22_introversion_extraversionWhere would you place yourself on the introversion to extraversion scale?Slider, 0 to 5Neither
q23_neurodivergence_impactHow much does neurodivergence (ADHD, autism, dyslexia, or similar) impact your daily work?Not applicable to me · Minimally · Moderately · SignificantlyNeither
q24_ageWhat is your age?Age in years; bucketed into cohorts only afterwardsNeither
q25_textAnything else you want to share?Free text, optionalNeither

Direction is the scale's direction, not its coding. It says which end of the answer is the better outcome, and nothing about reverse-scoring: the burnout composite reverses productive best, quality of life and plan completion when it combines them, and all three are listed here as higher is better, because they are.

Alpha if an item is removed (overall α = 0.797 over 595 complete cases)
Itemα without this itemVerdict
Overwhelm frequency0.761Keep
Priority difficulty0.759Keep
Overcommit frequency0.778Keep
External pressure needed0.772Keep
Replan frequency0.792Keep
Plan completion0.770Keep
Quality of life0.781Keep
Productive best0.783Keep
Pain-screener segments, collected at intake and not used in this edition
SegmentDefinition%Respondents95% CI (Wilson)
Screener: high (pain score ≥ 4)composite pain ≥ 457%343[53.0%, 60.9%]
Screener: mid (2–3)composite pain 2–338%229[34.2%, 42.0%]
Screener: low (< 2)composite pain < 25%30[3.5%, 7.0%]

The pain screener above was collected at intake to size an audience, and no figure in this edition uses it. The analysis runs on the burnout composite and the profiles it assigns.

Burnout bands, read off the profile rather than from a score cut-off (595 scored respondents)
BandProfiles it collectsRespondents
SevereFull Burnout75
StrugglingPlanning Chaos, Overwhelmed, Demand Overload272
At RiskDiffuse Strain, Productive but Paying, Quiet Struggle88
HealthyCoping, Thriving160

These bands are failure types read off the profile assignment, not rungs on a severity ladder: they name how execution breaks, not how badly. There is no score threshold behind them. The severity ladder is in section 2 of the report.

04 · ANALYSIS

Analysis

What was run
MethodHow it was used
CorrelationsPearson throughout, except drive against completion, which is reported as a Spearman rank correlation
Confidence intervalsbootstrap percentile, 2,000 resamples, for means
Wilson score intervalsfor proportions
Cohen's deffect size for differences between two group means
Multiple comparisonsnot corrected; this edition is exploratory and every figure should be read that way
Missing datacomplete-case per metric, no imputation, so a table's N can be lower than the sample
Data-quality checks (4 of 4 pass)
CheckPassValue
completenessLowest analysis column 89% (age); the optional closing free-text question is 52% and carries no figure
straightliners4 of 602 (0.7%) below SD 0.5 across 11 numeric items
reliabilityCronbach's α = 0.797, acceptable
sample-sizeN=602
What dropping the 4 flagged respondents would do to every headline mean
MetricAll respondentsFlagged removedShift
Productive best3.0653.064-0.001
Quality of life2.9772.973-0.004
Overwhelm frequency2.8362.831-0.004
Drive3.9473.9460.000
Burnout composite1.8881.884-0.004

The flagged respondents were kept. The table above is why the count is worth stating and not worth arguing about: on the pipeline's own convention it is four people, on the other common convention it is one, and dropping all four moves no published mean by more than 0.004.

Headline means with their confidence intervals
MetricScaleMean95% CI (bootstrap percentile)Respondents
Productive best0 to 5, higher is better3.06[2.99, 3.14]602
Quality of life0 to 5, higher is better2.98[2.89, 3.06]602
Overwhelm frequency1 to 5, higher is worse2.84[2.74, 2.93]602
Drive0 to 5, higher is better3.95[3.86, 4.03]602

Specific LLM extraction prompts, confidence thresholds, and profile-assignment cutoffs are proprietary methodology and not publicly released. Profiles are assigned by rule, not clustering. The aggregate data behind every figure in this report is published under CC-BY-NC 4.0, and the files are linked above.

05 · LIMITATIONS

Limitations

Known-groups check: the one contrast tested, and it did not separate
ContrastSplit atBelow the splitAt or aboveDifference (95% CI)Cohen’s dSeparates?
Role count and the burnout composite4 or more roles1.87 (392 people)1.92 (203 people)0.05 [-0.09, 0.19]0.06 (negligible)No

This is the one known-groups check we ran, and it did not discriminate. If the burnout composite measured strain from juggling roles, people carrying four or more roles should score higher than people carrying fewer; they score 0.05 higher on a 0-to-5 scale, and the interval on that difference straddles zero. Read it as a null result, not as validation. It is also a weak test of the composite: section 10 of the report publishes role count as a non-predictor of almost everything, so a null here is what that finding predicts. A contrast that should separate is the better test, and this edition does not publish one.

This edition has limits, and they are worth stating plainly. The sample is a convenience sample of high-agency working people, N=602, and it does not represent the general working population. Most of it was recruited through a paid participant panel against the screen set out above, so the people here are, by construction, in full-time work, inside an age band, doing most of their work on a computer, already using AI at work weekly, and running something of their own. Every figure should be read as describing people like that. The design is cross-sectional: every figure in the report is an association measured at a single point in time, and nothing here establishes cause. The instrument collected no physiological, sleep, time-tracking or buffer-time measures, and it followed nobody over time, so any claim that would need those is attributed external research or it is not made. Neurodivergence is one self-report item asking how much ADHD, autism, dyslexia or similar impacts daily work, answered on a four-point scale from not applicable to significantly. It measures impact, not diagnosis.

Three groups are small enough to name. The burnout taxonomy defines nine profiles; the ninth, Quiet Struggle, matched a single respondent and is excluded from every published figure. Productive but Paying holds 15 respondents, so its numbers are directional rather than precise. So is the planning inversion: only 41 respondents never plan at all, and the gap between them and the inconsistent planners does not reach statistical significance.

Two definitions carry more weight than their wording suggests. The severity groups in section 2 are cumulative rather than a partition, and each one is a threshold rather than a label. In crisis means a burnout composite of 3.0 or above. In distress means severe on any one of three measures: a burnout composite of 3.0 or above, a quality of life of 1 or below, or daily overwhelm. Under strain means a quality of life of 2 or less on the 0 to 5 scale, or overwhelm almost daily. Each group contains the one before it, so the shares do not sum. Thriving is the apex profile, outside crisis and distress, though two of its 37 respondents fall under strain. The pain screener in the segment-definitions table above was collected at intake and is not used in this edition’s analysis, which runs on the burnout composite and the profiles it assigns.

06 · PROFILE DETAIL

Profile detail

Per-profile item means. The first five are 1-to-5 frequency items where higher is worse; wellbeing priority is a 0-to-5 slider where higher is better
ProfileOverwhelmOvercommitPriority difficultyExternal pressureReplanWellbeing priority
The Replan Loop2.051.742.121.853.304.17
The Holding Pattern2.462.072.122.072.624.11
The Red Zone4.163.773.753.483.963.55
The Slow Burn3.432.763.002.833.214.12
The Volume Trap4.252.012.242.102.743.90
The Pressure Cooker2.262.872.112.892.424.15
The Benchmark1.621.241.271.271.624.35
The Hidden Cost3.132.202.272.333.133.33