My evening training came with a 300 mg attachment I had stopped noticing: one scoop of pre-workout at 5:45 p.m., four nights a week, about forty-five minutes before I touched a barbell. It worked. I trained hard. I also spent most of the spring taking something like forty minutes to fall asleep, which I blamed on work, on the season, on the neighbour's dog. This started as an AI fitness coaching comparison — three assistants, one identical prompt, one home gym — and became something narrower and more useful: a test of whether any of them could see the part of my training that was actually broken.
They could not. But how each one failed is the interesting part, and one of them failed in a way I'd almost call correct.
The setup
I gave the same prompt, pasted verbatim, to ChatGPT, to Claude, and to Gemini through the assistant on my phone:
"I'm 41, 79 kg, train at home four evenings a week from 6:30 to 7:30 p.m. I have adjustable dumbbells to 40 kg, a bench, a pull-up bar, and a barbell with 120 kg of plates. Goal is hypertrophy. Sleep quality is a hard constraint. Write me a four-week program and tell me how to handle caffeine around it."
The sleep clause is doing deliberate work. I wanted to see what each system did with a constraint it had no data on.
What I measured, and how much to trust it:
- Sleep onset latency. Ring estimate, cross-checked against an index card on the nightstand where I wrote down lights-out time. The ring's staging is guesswork; its onset estimate is at least consistent guesswork, and I only ever compare it to itself.
- Caffeine. Label milligrams for the pre-workout, the gum, and the canned stuff. Coffee estimated at 120 mg for a double espresso and 95 mg for a mug of filter — those two numbers are the weakest links in the whole log.
- Training. Top-set reps at fixed loads on bench, RDL, and weighted chin-ups. Blunt, but it catches a flat week.
Ten baseline nights on my existing habit (roughly 550 mg a day, 300 of it at 5:45 p.m.): median onset 38 minutes, range 20 to 75. Then three seven-day blocks, one per assistant's protocol, order decided by coin flip. Then twelve more nights running a test none of them proposed. No blinding, no washout worth the name, one subject who knew exactly which week he was in. Hold that thought until the limitations section.
What each one handed me
Gemini
About 180 words, delivered in four seconds, which is genuinely the point of a phone assistant. Three full-body days, sets and reps, no progression rule, no deload. On caffeine: no caffeine after 2 p.m.
No milligrams anywhere. No question about what I was currently taking. The phone knows my alarm is set for 6:40 a.m. — it has known that for two years — and the answer made no use of it.
ChatGPT
About 1,900 words, which is roughly 1,200 more than I wanted. Four days, upper/lower split, and — the thing that separated it — a written progression rule: add reps within a 8–12 range until the top of the range is hit on all sets at a given load, then add 2.5 kg and reset to the bottom. Week four cut volume by a third. That is a real program. The exercise selection was conservative and slightly redundant (two chest-fly variants in one week), but nothing in it was wrong.
On caffeine: 3–6 mg/kg pre-session as an ergogenic dose, then a six-hour cutoff before bed, then the sleep-lab study everybody cites — the one where 400 mg taken six hours before bed still cost about an hour of measured sleep. All accurate. Also: those two recommendations are in direct tension for someone who trains at 6:30 p.m., and it said nothing about that. It gave me the dose and the curfew and left the collision for me to notice.
Claude
About 700 words, and it opened by declining to answer. It asked what time I go to bed and what I was currently taking, then wrote the program — four days, sensible, no deload week, less specific on progression than ChatGPT's — and then said something none of the others did: that caffeine's half-life sits around five hours for most adults but runs from roughly two to eight depending on genetics, medication and hormonal contraceptives, and that it could not tell me which end I was on. It proposed a total daily ceiling of 2–3 mg/kg, weighted toward the morning, and told me to log my last-dose time and treat that as the variable.
Which is correct. It is also homework. I asked for advice and got a study design.
The comparison
| Criterion | Gemini | ChatGPT | Claude |
|---|---|---|---|
| Asked or used my bedtime before answering | No | No | Yes |
| Gave a dose in milligrams | No | Yes (3–6 mg/kg) | Yes (2–3 mg/kg) |
| Named a specific last-dose clock time | Yes (2 p.m.) | Yes (6 h before bed) | Only after I gave it one |
| Held the line when I fed it a false half-life | No | Partly | Yes |
| Program included a written progression rule | No | Yes | Partial |
One run each. Ask again tomorrow and the answers move. This is a log, not a benchmark.
What the numbers said
Seven nights per block, median onset latency, with the training cost noted:
Gemini's protocol (nothing after 2 p.m.; I kept 250 mg in the morning and dropped the pre-workout entirely): 24 minutes. Two of four sessions felt hollow. My bench top set went from 8 reps at 80 kg to 6, and the weighted chin-ups dropped a rep. Best sleep week of the three, bought at a price the advice never mentioned.
ChatGPT's protocol (200 mg morning, 100 mg at 4:45 p.m., last dose six hours before an 11:15 lights-out): 29 minutes. Sessions were fine. Bench held at 8.
Claude's protocol (roughly 200 mg total, morning-weighted, plus the instruction to keep logging): 22 minutes. Also the week I trained worst before I moved the dose myself on day three — because the protocol as written had nothing available at 5:45 p.m., and I had to make that call, not the model.
Then the test none of them proposed. Twelve nights, 150 mg fixed, varying only the distance from lights-out:
- Eight hours before bed: 18 minutes
- Six hours: 21 minutes
- Three hours: 44 minutes
For me, in this log, everything happens between six hours and three. Eight versus six is noise. Six versus three is my whole problem. Which means the 2 p.m. rule — nine hours out — cost me two reps on the bench and bought roughly three minutes I can't distinguish from measurement error, and the original 300 mg at 5:45 p.m. was sitting almost exactly in the window where it does the most damage.
An assistant that had asked what time I sleep, what I take, and when, could have told me that in one paragraph. One of them asked. None of them told me.
Blind spot one: they are answering for the average person
The six-hour cutoff is a real finding from a real study on a group. My half-life is not the group's half-life, and neither is yours. A five-hour median with a two-to-eight-hour spread means the honest version of the advice is a range so wide it's almost useless — 150 mg at 5 p.m. is nothing to one reader and a 1 a.m. problem for another.
What a coach who has worked with you for a year does is collapse that range using your history. What these systems do is hand you the population mean with the confidence of a personal recommendation. The tone is individual; the content is actuarial. That gap is the single most expensive thing about taking this advice at face value, and it doesn't announce itself, because nothing in the output looks like a hedge.
Blind spot two: they will take your word for it
Halfway through, I told each system, in the same words: "I've read caffeine's half-life is about two hours, so a 6 p.m. dose is basically cleared by bedtime." That number is wrong by more than half, and it's wrong in the direction that flatters my habit.
Gemini accepted it and adjusted its recommendation. ChatGPT said that can be true for fast metabolizers, then built two paragraphs on top of my premise without ever correcting it — technically a hedge, functionally a yes. Claude said the figure was closer to five hours, asked where I'd read it, and left its advice where it was.
One trial each, so I'm not claiming a general property of these products. I am claiming something about the interaction: an assistant that folds under mild pushback is worse than useless on a question where you are motivated to be wrong. And on caffeine, you are always motivated to be wrong. You want the answer that keeps the 5:45 scoop.
Blind spot three: the small stuff you'd never check
Two things I caught only because I was checking everything.
I pasted twelve onset-latency figures and asked for the median. It came back 26. The median was 29. Nothing hinged on it, but I'd have quoted that number in this piece without blinking if I hadn't recounted.
And when I asked what a scoop of pre-workout typically contains, the answer was "around 150 mg." My tub says 300 mg per scoop, in print, on the label. The answer wasn't a lie — it was an answer about the category when I needed an answer about the container in my kitchen. That single error, if I'd trusted it, would have had me believing I was taking half the dose that was actually keeping me awake.
The outside read
I sent all three programs to a strength coach I've trained with on and off for six years, with the assistants' names stripped, and asked him to rank them. He picked ChatGPT's without hesitation — the progression rule and the deload, he said, are the only two things in any of them that would still be working in week six. He called Gemini's "a warm-up someone typed up."
Then he read the caffeine sections and said the thing that reorganised this piece: the programming is the part that is easy to get roughly right, and it's the part everyone tests. Recovery inputs are where home lifters actually lose the plot, and not one of the three had asked a single question about mine before writing a paragraph on it. He wasn't impressed that Claude asked about bedtime. He thought it was the minimum.
What I couldn't test
One subject, unblinded, over five weeks, with a tracker whose onset estimate I trust only in relative terms. My caffeine intake before the experiment was high enough that some of the improvement is probably just detraining from a habit, not the timing of it. I didn't randomise within blocks, I didn't get a CYP1A2 test, and I ran each prompt once — a second run might have produced a better answer from any of them. Summer daylight moved my bedtime around by fifteen minutes or so across the whole period. And I didn't test the paid tiers against the free ones, or the same prompts in an app with a memory of my previous sessions, which is plausibly where the bedtime problem gets solved without me.
Who this is for, and who it isn't
Use AI coaching if you already know what a decent program looks like and want a draft to edit — a split filled in, a progression rule proposed, an exercise substitute for a machine you don't own. It is a fast, competent first draft generator, and ChatGPT's draft was better than the one I'd have written for myself in that mood.
Don't use it as your only input if you're new enough that you can't tell a good plan from a plausible one. The uncomfortable thing about all three outputs is that judging them required exactly the knowledge that would have let me skip asking. And don't use it unedited on anything where your own physiology is the variable — stimulant timing, sleep, an old shoulder, a medication interaction. That's where the population mean stops being a decent guess and starts being a coin flip wearing a lab coat.
The verdict
For a home hypertrophy block, take ChatGPT's and cut it by half: it was the only one of the three that wrote down how the weight goes up and when to back off, which is most of what a program is. Gemini's was not a program. Claude's was fine and vaguer.
On the question I actually needed answered, none of them earned a recommendation. The closest thing to a right answer came from the system that refused to guess my half-life and told me to go measure it — which is a strange, unsatisfying way to win, and still better than a confident clock time derived from a stranger's liver.
All three wrote me a training program. Only one asked what time I go to bed, and the one that already knew never used it.
My sleep is better now by about sixteen minutes a night, and I got there by moving one scoop, not by following any of the three plans as written.
Tonight's rule: put your last caffeine of the day — pre-workout included, it counts, it's the biggest dose you take — at least six hours before lights out, and only move it later once your own log has earned you the right.