The Evidence Audit
What I Can Prove, What I Can't, and What Time Teaches That Labs Can't
The problem this addresses
Can I trust what my coach is telling me?
The Problem
In 2015, a group of researchers tried to reproduce 100 published psychology studies. These weren't fringe findings. They were peer-reviewed, statistically significant results from respected journals. The kind of research that shapes TED Talks, bestsellers, and the frameworks your coach uses.
Thirty-six percent replicated.
Nearly two-thirds of published findings couldn't be reproduced when someone else tried. Power posing? The hormonal effects vanished. Ego depletion? Twenty-three labs, 2,000 participants, effect size of approximately zero. The study that "proved" priming people with "old" words made them walk slower? Gone.
I'm a coach. My livelihood depends on my frameworks working. That makes me exactly the kind of person you should be skeptical of on this point. So here's what I did: I audited every major source in my toolkit against two standards. The first was the peer-reviewed evidence. The actual studies, the meta-analyses, the replication attempts, the effect sizes. The second was older and harder to argue with: time. Has this practice survived centuries of continuous use across independent cultures? Or is it a recent invention dressed up as science?
Some of what I found confirmed what I use. Some of it didn't. Some of it revealed problems the coaching industry refuses to discuss. This document is the result.
Why This Document Exists
The coaching industry has a credibility problem it doesn't want to talk about. It also has a harm problem it doesn't want to talk about. And it has a measurement problem that makes both worse.
The credibility problem. The headline ROI figure you'll see everywhere, "788% return on investment from executive coaching," traces to a single case study of 43 managers at one company, commissioned by a coaching vendor, where executives were asked to self-estimate the business value of their coaching. Not peer-reviewed. No control group. The participants were asked, "How much value do you think coaching created for you?" and their answers were monetized. That number has been cited thousands of times across the industry as though it were established science. It wouldn't survive a first-year research methods class.
The harm problem. Sixty-seven percent of coaching clients report at least one negative effect. Averaging 3.46 negative effects per client. Reduced job satisfaction in 31%. Triggering of psychological problems the coach couldn't handle in 26%. Decreased sense of meaning toward work in 17%. Those numbers don't appear in any coaching marketing material. They don't appear in the meta-analyses that produce the industry's positive headlines. They exist in a separate literature that nobody is connecting.
The measurement problem. The best peer-reviewed meta-analysis of coaching (de Haan & Nilsson, 2023) analysed 37 randomised controlled trials. The overall effect size was g = 0.59. Moderate. Real. But look underneath that number. The prediction interval was -0.20 to 1.38. That means the true effect could be anywhere from slightly negative to very large. The individual study effects ranged from 0.02 to 1.98. When you correct for publication bias, the estimate roughly halves to g = 0.27. When someone other than the coached person measures the change, the effect halves again.
The average is the temperature of a room that's on fire in one corner and freezing in another. The average is comfortable. The reality is not.
Here is what I believe: coaching works for some people, in some contexts, some of the time. The evidence supports a moderate positive effect on self-efficacy, goal attainment, and wellbeing. But the industry consistently overpromises, ignores the harm data, and presents averages that hide more than they reveal. Most coaches couldn't tell you the effect size of their own approach, the contraindications of their own techniques, or the base rate at which their clients would have improved without them.
I'd rather you trust me because I've done the work than because I've quoted an inflated statistic.
What Time Teaches That Labs Can't
Before getting to specific frameworks, a word about what counts as evidence.
A 12-week randomised controlled trial with 50 university students is one form of evidence. It has precise effect sizes and confidence intervals. It also tests a narrow intervention, on a narrow population, under artificial conditions, measured with instruments that may not capture what actually matters.
A practice that has survived 2,000 years of continuous use across four or five independent civilisations is a different form of evidence. It has no effect sizes. It also passed a test no lab can replicate: billions of human lifetimes, across every conceivable condition, with the harshest possible selection pressure. Practices that didn't work didn't survive. The ones that survived are the ones that helped people enough to keep getting passed down, generation after generation, century after century, culture after culture.
I've started weighing both. Where a practice has both lab evidence and civilisational track record, I trust it most. Where it has only lab evidence, I hold it more lightly. Where it has only a track record, I pay attention to the track record. Where it has neither, it's gone.
The practices that anchor my work are old. Breathwork is older than psychology. Socratic questioning is older than psychotherapy. Reflective writing is older than the journal that published Pennebaker's study on it. Structured accountability through community is older than the coaching industry by roughly 200,000 years.
Modern research confirms these practices. It didn't create them.
What I Use and What the Evidence Says
The Lindy Stack (Strong Ground)
These practices have both peer-reviewed support AND centuries or millennia of continuous use across independent traditions. They're the safest ground I stand on. Not because of any single study. Because time already ran the experiment.
Breathwork. Pranayama has two thousand years of continuous practice across Hindu, Buddhist, Taoist, and Christian hesychast traditions. Four independent civilisations discovered that controlling your breathing controls your mind. Modern science drew the wiring diagram: controlled breathing activates the vagus nerve, shifting the autonomic nervous system toward parasympathetic dominance. The Wim Hof endotoxin study (Kox et al., 2014, published in PNAS) showed trained volunteers could voluntarily modulate their immune response. Stanford's cyclic sighing study (2023, 108 participants) found five minutes of daily structured breathing outperformed mindfulness meditation for mood improvement. When I use breathing in the Co-Regulation Protocol, I'm using a tool that monks, soldiers, and philosophers trusted long before neuroscientists explained why.
Cognitive reframing (Socratic questioning with a modern name). Socrates developed the method 2,400 years ago, describing it as midwifery -- helping people give birth to their own understanding. Epictetus wrote in the 1st century: "Men are disturbed not by things, but by the views which they take of things." That sentence is, word for word, the foundational premise of Cognitive Behavioural Therapy. Aaron Beck explicitly credited the Stoics when developing CBT. Albert Ellis stated that his approach's central principles were "originally discovered and stated" by them. Meta-analyses give CBT effect sizes of g = 0.71 for depression and g = 0.56 for anxiety. When I help a founder reframe a result from "I failed" to "that's data," I'm not using a modern technique. I'm using the oldest therapeutic method in Western civilisation, validated by the largest evidence base in modern psychotherapy.
Reflective writing. The Stoic practice of hypomnema -- personal philosophical notebooks -- dates to at least the 1st century BCE. Seneca practised nightly self-examination: "Which of your ills did you heal today? Which vice did you resist?" Marcus Aurelius wrote the Meditations during military campaigns on the Danube frontier around 170 CE. Never intended for publication. Still in print 1,850 years later. Pennebaker's modern research (400+ studies, d = 0.16) confirms what every serious practitioner of self-examination already knew: writing clarifies thinking, and clarity reduces suffering. A 2022 network meta-analysis found enhanced expressive writing was as effective as traditional psychotherapy for trauma processing. But here's a caveat: it doesn't work for everyone. People with low emotional expressivity showed increased anxiety. People with active PTSD and rumination tendencies found writing gave negative thoughts more emphasis. When I use the Pennebaker letter in the Forgiveness Protocol, I match it to the client. It's not a blanket prescription.
Affect labelling ("The story I'm making up"). Lieberman et al. (2007) at UCLA demonstrated that putting feelings into words reduces amygdala activation and increases prefrontal cortex activity. Naming an emotion dampens its intensity through a measurable neural pathway. This also has a Lindy pedigree: the entire contemplative tradition across Buddhism, Stoicism, and monastic Christianity involves naming internal states as a practice. When I ask a founder to articulate the fear out loud, the processing isn't metaphorical. It's neurological.
Identity-level change as the deepest change. Oyserman's Identity-Based Motivation framework, Self-Determination Theory's integration research, and addiction recovery literature all converge: when behaviour becomes part of who someone sees themselves as, it sticks without requiring willpower. This insight is newer in its academic form, but every initiation ritual, every military training programme, every religious conversion practice in human history understood that identity transformation is the mechanism of lasting change. When the Anti-Fragile Founder focuses on decoupling identity from outcomes, it's drawing on both the modern evidence and something far older.
Structured accountability through relationship. This is probably the oldest intervention in human existence. Tribal structures, religious communities, guild systems, military units, the master-apprentice tradition. Humans are social primates who regulate behaviour through belonging and mutual obligation. The coaching meta-analyses confirm what these traditions always demonstrated: the working alliance (the quality of the coach-client relationship) correlates with outcomes at r = .41 across 27 samples. That correlation is larger than the effect of any specific coaching methodology. The research calls this "common factors." The ancient world called it mentorship, spiritual direction, and philosophical friendship.
The Research-Validated Stack (Strong Ground)
These have strong peer-reviewed support but shorter track records. Validated in labs, not yet validated by centuries.
Anchoring (Kahneman). The Many Labs project tested anchoring across 36 labs and thousands of participants. It replicated consistently with large effect sizes (d = 0.5 to 1.0+). Independent labs find it regardless of who runs the study. The mechanism is straightforward: you give someone a number, they adjust insufficiently from it. This shows up in courtroom sentencing, real estate pricing, salary negotiations.
Shame versus guilt (Tangney, Brown). A meta-analysis of 86 studies found shame-proneness correlates with depression at r = .43. Guilt-proneness correlates with prosocial behaviour. The distinction, "I am bad" versus "I did something bad," has been validated across populations for decades.
Vulnerability enables connection (Aron, Brown). Aron's "Fast Friends" procedure showed that escalating reciprocal self-disclosure generates interpersonal closeness in under 45 minutes. Replicated by Sprecher in 2021. Attachment theory provides the developmental framework. A caveat: this research was conducted between people with equal power and no professional stakes. Vulnerability in a boardroom, an investor meeting, or a competitive leadership environment carries different risks. Before I coach anyone toward more vulnerability, I assess whether their environment is actually safe for it.
The vulnerability paradox. Bruk et al. (2018) tested this in the Journal of Personality and Social Psychology across seven experiments. We rate others' vulnerability as courage and our own as weakness. Every time. The finding is clean, the journal is top-tier, and it replicates.
Perfectionism as distinct from high standards (Hewitt & Flett). A 2023 meta-analysis: perfectionistic concerns correlate with depression and anxiety at r = .38 to .43. Perfectionistic strivings show much smaller correlations at r = .10 to .21. The distinction shows up in measurable outcomes. Perfectionism has also increased generationally across a meta-analysis of 41,641 participants. Young people increasingly feel others demand perfection of them.
Self-efficacy as the primary mechanism of coaching. This may be the most important finding in this entire audit. Across coaching meta-analyses, self-efficacy and goal attainment show the largest effect sizes (g = 0.59 to 1.29). Coaching creates mastery experiences through structured goal-setting. Small wins build confidence. Confidence produces action. This maps directly onto Bandura's four sources of self-efficacy, and it explains what coaching actually does better than most coaching frameworks explain what coaching does.
The Moderate Ground
These concepts have support but carry caveats. The direction is right. The magnitude is often overstated.
Gratitude practice. A 2025 mega-analysis in PNAS (145 papers, 24,804 participants, 28 countries) found a small effect: Hedges' g = 0.19. When compared to doing something else rather than doing nothing, the effect dropped to essentially zero. Gratitude journaling is a useful tool. It's not the foundation of a philosophy.
Mindfulness and attention training. Amishi Jha's research shows mindfulness protects against attention degradation during high-stress periods. The evidence is real but modest, and a 2016 analysis found 88% of published mindfulness RCTs reported positive results against an expected 66%. Publication bias is inflating the field. Mindfulness also has a harm profile that most proponents don't discuss: 83% of participants in one study reported at least one side effect, 37% reported negative impacts on functioning, and 6-14% reported lasting adverse effects including dissociation and hyperarousal. A researcher who studies meditation harm estimates 30-50% of meditators experience an adverse event. Mindfulness is maintenance, not magic. And it has contraindications.
Coaching itself. The honest effect size is g = 0.5 to 0.6 before correcting for publication bias. After correction: roughly g = 0.27. Self-reported outcomes produce effects approximately twice as large as outcomes measured by people around the coached individual. The meta-analyses found no significant moderation by coaching duration. Longer engagements did not produce significantly larger effects than shorter ones.
What I've Removed
NLP (Neuro-Linguistic Programming). Three major systematic reviews converge. Classified as pseudoscience. I don't use it.
Ego depletion. Willpower as a finite resource that depletes like a gas tank. Twenty-three labs. Over 2,000 participants. Effect size of approximately zero.
Power posing. The hormonal effects couldn't be replicated.
The alkaline diet mechanism. Physiologically impossible. The man who popularised it was convicted three times for practising medicine without a licence and sentenced to prison.
The "788% ROI" and similar coaching statistics. Self-reported estimates collected by organisations with direct financial interest.
Joe Dispenza. Claims about changing DNA through meditation. No peer-reviewed publications supporting the core claims.
Growth mindset as a framework. Carol Dweck's research shows an average effect of d = 0.08. Three percentile points. The tutor matters more than the mindset. When I talk about identity and adaptability, I draw on the identity-based motivation research, which has a much stronger evidence base.
The Number Behind the Number
This section exists because the coaching industry reports averages. Averages hide more than they reveal.
The g = 0.59 coaching effect looks moderate and reassuring. Here's what's underneath it.
The range is enormous. Individual study effects go from 0.02 to 1.98. That's not a bell curve around a moderate centre. It's the signature of a domain where coaching either does almost nothing or produces dramatic change, depending on something the average doesn't capture.
The prediction interval crosses zero. If you ran a new coaching study tomorrow, the best statistical prediction based on existing evidence is the true effect could be anywhere from slightly negative to very large. The average tells you almost nothing about what to expect from any specific engagement.
Most of the measured benefit may come from a minority of clients. In psychotherapy (the closest parallel), more than half of patients receiving therapy did not respond at all. 41% responded versus 16-17% in controls. The number needed to treat is roughly 5 -- you need to coach about 5 people for 1 additional person to benefit beyond natural improvement. No coaching study has calculated this number. The field doesn't think in these terms yet.
Nobody has looked at the distribution shape. No individual-level outcome analysis. No variance ratio analysis. No distributional fit test. We don't know whether coaching outcomes follow a normal distribution, a fat-tailed distribution, or something else entirely. The field reports the mean. The full distribution, the thing that would actually tell you your odds, doesn't exist.
What this means for you: I can't tell you coaching will produce a specific outcome. I can tell you that the structure I provide (focused attention, structured goals, honest accountability, frameworks tested against both evidence and time) is the best version of what the research says works. Whether it works for you specifically depends on variables that no average can predict.
The Graveyard
Most coaches will never show you this section. The coaching industry collects evidence from the people who succeeded and calls it proof. The people who didn't succeed don't write testimonials. They don't appear on stage. They blame themselves. They disappear.
Here's what the research says about when coaching causes harm.
67.6% of coaching clients reported at least one negative effect. Averaging 3.46 negative effects per client. Most were low-to-medium intensity. But the sheer frequency is striking: two-thirds of people coached experience something negative. Specific harms: reduced job satisfaction (31%), triggering psychological problems the coach couldn't handle (26%), decreased sense of meaning toward work (17%). Higher relationship quality and coach expertise predicted fewer negative effects. The absence of supervision predicted more.
Identity-level interventions carry real risk. Destabilising someone's sense of self is not a neutral act. The clinical literature is explicit: destabilisation is a recognised technique in both therapy and brainwashing. The difference is supervision, consent, dosing, and the capacity to reconstruct. When a trained therapist carefully destabilises a rigid pattern within a supervised clinical relationship, that's therapy. When an unqualified coach destabilises someone's identity during a high-stress period without clinical training, without supervision, without the ability to manage the fallout, that's a risk vector. I take this seriously. After any identity-level intervention, I follow up within 48 hours.
Expressive writing has contraindications. People with low emotional expressivity showed increased anxiety after writing. People with active PTSD and rumination tendencies found writing amplified negative thoughts. When distress is intense or maintained by rumination, writing exercises are "not advisable." I don't use the Pennebaker letter as a blanket intervention.
Vulnerability coaching in unsafe environments can backfire. The research on vulnerability was conducted between people with equal power and no competitive stakes. Founders operate in environments where vulnerability has real political and financial costs. Before I coach anyone toward more openness, I assess: is this environment actually safe for that? What are the power dynamics? Some boardrooms will punish clarity.
The coaching industry is unregulated. No licensing. No educational prerequisites. No supervised practice requirements. An ICF Associate Certified Coach requires 60 hours of training. A clinical psychologist requires roughly 10,000. A 170:1 ratio. A ProPublica investigation found that about a third of Utah therapists whose licenses had been revoked continued working as "life coaches." One-third of 3,500 licensed therapists surveyed reported clients who'd been harmed by a life coach. The testimony: coaches who had clients "deep dive into their trauma, which sent them into an emotional spiral and then did not provide them with any skills to cope." Four of five ended up hospitalised with severe suicidal ideation.
I'm not a therapist. I don't treat clinical conditions. When I recognise that a client needs clinical support, I refer. That boundary isn't optional.
Large group events have a documented body count. Lifespring: more than 30 lawsuits, at least four documented deaths. James Arthur Ray: three deaths at a 2009 sweat lodge, 18 injured, $10,000 per ticket, convicted of negligent homicide. Multiple event organisations have documented psychotic episodes, hospitalisations, and post-event suicides. I don't run large group interventions. I work with individuals and small groups where I can see what's happening and respond to it.
What Coaching Actually Does (The Honest Version)
The research paints a specific picture, and it's not the picture the industry paints.
The common factors are the headline, not the footnote. Research consistently shows coaching effectiveness is not driven by specific methodology. It's driven by common factors: relationship quality, empathic understanding, goal focus, accountability, positive expectations. The working alliance predicts outcomes more strongly than any coaching model, framework, or technique.
This means my protocols -- the Decision Architecture, the Optionality Engine, the Co-Regulation Protocol -- are frameworks for delivering structured attention. They organise the conversation. They create scaffolding for the relationship. They're not the active ingredient.
The active ingredient is older and simpler: someone with relevant experience paying sustained, focused attention to your goals, your patterns, and your growth. Socrates did this. The Stoic philosophers did this. The master-apprentice tradition did this. Every religious community with a spiritual direction practice did this. The coaching industry reinvented a 2,400-year-old relationship and gave it a billing code.
That's not a diminishment of what I do. It's a clarification. My value is not my framework library. My value is the quality of attention I bring, the honesty I maintain, and the accountability I hold. The frameworks help me do those things well. They're the scaffolding, not the bridge.
The rooster problem. A rooster crows every morning. The sun rises every morning. The rooster concludes he causes the sunrise.
I work with founders who are already motivated, resourced, and growth-oriented. They were selected for drive before I ever met them. Some percentage of their improvement was coming anyway. I can't tell you the exact percentage. Neither can the research. The self-selection problem in coaching studies has never been adequately controlled for.
What I can do is be honest about it. I track what happens when clients end engagements. I ask the uncomfortable questions: what would you have figured out on your own? What did coaching give you that you genuinely could not have gotten any other way? When the honest answer is "accountability and someone to think out loud with," that's real value. It's just not the transformative-guru value the industry likes to sell.
What coaching can do: Build self-efficacy through structured small wins. Hold focus on what matters when distraction is everywhere. Maintain accountability that a founder's environment doesn't provide. Offer pattern recognition from hundreds of engagements. Challenge in a way that's direct without being destructive. Name the thing nobody else will say.
What coaching cannot do: Guarantee business outcomes. Replace therapy for clinical conditions. Substitute for the skill, preparation, timing, and luck that determine business results. Override the structural and environmental factors that shape most of what happens to a company.
The Filter I Use (And You Can Too)
If you want to evaluate any psychological claim, whether it comes from me, a book, a TED Talk, or someone on the internet with a PhD, run it through these six questions:
Has it survived? How old is the practice? Has it been used continuously across multiple independent cultures? A practice that four independent civilisations discovered independently has passed a test that no 12-week study with 50 undergraduates can match. Time is the ultimate sample size.
Has it been independently replicated? Not "cited a lot." Not "published in a good journal." Has a different lab, with different researchers who didn't invent the concept, run the study and found the same thing?
What's the effect size -- and what's the distribution? An average of 0.5 could mean everyone gets a small benefit. It could also mean a few people get dramatic results while most get nothing. Ask about the range, not just the centre.
What are the contraindications? Every medication lists side effects. Every surgical procedure has a known complication rate. If someone is selling you a psychological intervention and can't name a single person it might harm, they haven't looked.
Does the researcher talk about limitations? "In some contexts, for some people, with adequate support" is honest science. "This changes everything" is marketing.
Who funded the research? The coaching industry funds most coaching research. When a trade body studies its own effectiveness, the structural conflict is obvious.
Where This Came From
This audit draws on three tiers of evidence.
The first is peer-reviewed research: meta-analyses, systematic reviews, and randomised controlled trials across psychology, neuroscience, organisational behaviour, and clinical research. Key sources include the Open Science Collaboration (2015) replication study, de Haan & Nilsson's 2023 coaching meta-analysis, Tangney's shame research, Bandura's self-efficacy framework, Lieberman's affect labelling research, Kox et al.'s breathwork study, the Many Labs replication project, and Schermuly's research on negative coaching effects.
The second is the Lindy test: civilisational track record. Practices that survived continuous use across independent traditions for centuries or millennia, including Socratic questioning, pranayama, contemplative practice, reflective writing, and structured community accountability.
The third is the graveyard: documented harm. Schermuly and Grassmann's side-effects research, Britton's meditation adverse effects studies, Lieberman's encounter group casualties study, ProPublica's coaching harm investigation, and the documented body count from large group awareness training.
I'm aware this is unusual for a coaching practice. Most coaches cite the 788% ROI and move on. I'd rather show you the full picture and let you decide.
You didn't hire me to be comfortable. You hired me to be honest. And the most honest thing I can tell you is this: the practices that anchor my work are old. The relationship that delivers them is the mechanism. The frameworks organise the conversation. The outcomes are real but not guaranteed. And I track the failures, not just the wins.
The coaching industry that can't name its harm rate, can't describe its outcome distribution, and can't tell you what would have happened without the coaching is an industry that hasn't done the work.
I've done the work. This is what I found.
These protocols work on their own.
They work differently with someone in the room.
Interactive Tool
Where does your time actually go?
The Time Audit shows you the gap between where you think your time goes and where it really goes. Two minutes to the truth.
Start your Time Audit