Are AI Calorie Trackers Accurate? What the Gold-Standard Test Found
In the only kind of study that can settle this, a food-recognition app was compared against
doubly labelled water, the gold standard for measuring energy in free-living people.
The app estimated 1,905 kcal/day against a measured 2,235 kcal/day,
a bias of −330 kcal/day. The number that matters more is the spread: the limits of
agreement ran from −1,504 to +845 kcal/day. On a given day the estimate could be
1,500 calories under or 845 over. A systematic review of 52 papers puts relative error for calories
anywhere between 0.10% and 38.3%, and concludes the tools need more development
before deployment as stand-alone dietary assessment methods. Any page advertising ±1% accuracy is
describing something the published literature has never measured.
This question has a real answer, and it is more interesting than either “AI is amazing” or “AI is
useless.” The research exists, it is specific, and almost none of it appears in the pages ranking for
this question.
The test that actually settles it
Most app-accuracy claims compare an app’s estimate against a nutrition database. That tells you
whether the app matched a table. It does not tell you whether either one matched what the person ate.
Doubly labelled water is different. It measures energy through isotope tracking in
the body over days of ordinary life, and it is the reference method against which other dietary
assessment tools are judged.
A study in Frontiers in Nutrition (2023, PMID 37810925) ran that comparison. Thirty adult
women of normal body weight, seven days, free-living conditions, app estimates against DLW:
- App estimate: 1,905.5 kcal/day (SD ± 531.1)
- Measured by DLW: 2,235.2 kcal/day (SD ± 456.5)
- Bias: −329.6 kcal/day
- Limits of agreement: −1,503.8 to +844.5 kcal/day
The researchers also found no significant linear relationship between the two (R² = 27%,
p = 0.50).
Read the second-to-last line again
The bias of −330 kcal/day is the headline, and it is not the important number.
The limits of agreement span 2,348 calories.
That is the range within which an individual day’s estimate is expected to fall. Under by 1,500, or
over by 845, and both are inside the expected behaviour of the tool. For someone eating around 2,000
calories, that is a band wider than their entire day’s intake.
A consistent bias is manageable. If an app were reliably 330 calories low, you could work with the
trend and ignore the absolute number. A 2,348-calorie spread is a different problem, because it means
today’s figure carries very little information about today.
It still beat the traditional method
This is the part that stops the article being a hit piece.
The same study measured 24-hour dietary recall, the interview method used in nutrition research for
decades, against the same DLW reference. Its bias was −543.0 kcal/day, with limits of
agreement from −1,802.5 to +716.5.
Worse on both counts. The authors concluded that the app provides a closer representation of energy
intake in adult women with normal body weight than 24-hour recall does, while adding that further
research is needed to determine clinical relevance.
So the honest summary is not that AI food tracking is bad. It is that measuring what a human
being eats is genuinely hard, every available method is wide, and the AI version is currently
the least wide of the practical options.
The 52-paper picture
A systematic review in Annals of Medicine (2023, PMID 38060823) covered 52 papers published
between 2010 and 2023, with convolutional neural networks used in 79% of them.
Relative error for calories: 0.10% to 38.3%. For volume: 0.09% to 33%.
That range is the finding. Under ideal conditions, with a single simple food photographed clearly, AI
estimation can be extremely close. With a mixed plate, sauces, layered dishes or an awkward angle, it
can be off by more than a third.
The authors’ own conclusion is the sentence to remember: the tools currently available need
more development before deployment as stand-alone dietary assessment methods in nutrition
research or clinical practice.
The methodological detail almost nobody mentions
In that review, the ground truth used was:
- Calculation using nutrient tables: 51%
- Weighed food: 27%
So in roughly half the studies, “accuracy” means the AI agreed with a database, not with a scale.
That matters because the database is itself an approximation. A restaurant portion of pasta is not the
database’s portion of pasta. An app can score well against a nutrient table and still be some distance
from the plate in front of you, and a study built that way cannot detect the difference.
Which foods break it
A 2024 study in Nutrients (PMID 39125452) assessed 18 nutrition apps in depth, drawn from 800
screened, comparing outputs against food records for Western, Asian and Australian-guideline diets.
Two apps came out highest for accuracy: MyFitnessPal at 97% and
Fastic at 92%.
The more useful result was directional. Manual logging apps overestimated energy for Western
diets by an average of 1,040 kJ and underestimated Asian diets by an average of
−1,520 kJ.
Same apps, opposite errors, decided by what was on the plate. Mixed dishes, shared plates, sauces and
foods poorly represented in Western-built databases are where the estimates go wrong, and they go wrong in
a predictable direction rather than randomly.
About those ±1% benchmark claims
Search this topic and you will find sites publishing tables with figures like ±0.9%, ±1.0% or ±1.7%
portion estimation error, attributed to independent testing of thousands of lab-weighed meals.
Set those against the published literature: a range of 0.10% to 38.3% across 52 peer-reviewed papers,
and limits of agreement spanning 2,348 calories against the gold-standard method.
Both cannot be true.
Three things are worth checking on any such page before you believe its table. First, whether its
numbers are internally consistent, because the same app is currently credited with different figures on
different pages of the same claim ecosystem. Second, whether different “independent benchmarks” crown
different winners, which is what you would expect if each site had a preferred product rather than a
method. Third, the site’s own claims about itself: at least one such benchmark describes itself as
established in 1999, which predates by two decades the technology it says it tests.
None of that proves any specific figure is wrong. It does mean the burden of proof sits with the page
claiming a precision that peer review has never achieved.
What accuracy do you actually need?
This is the question the whole argument skips, and the answer depends on what you are doing.
For tracking a trend, consistency matters more than accuracy. If the tool is
reliably low by a similar amount, the direction of travel over weeks is still readable. That is a real
use, and it is the one most people actually have.
For a specific numeric target, such as hitting an exact deficit or matching insulin
to carbohydrate, the spread in these studies is too wide to rely on. That is why the carbohydrate-counting
literature reports mean absolute errors in the range of 13 to 20 grams per meal for AI tools, which is a
clinically meaningful amount for someone dosing on it.
For adherence, accuracy may be the wrong metric entirely. Self-monitoring is one of
the better-supported behavioural components of weight management, and a tool that is imprecise but gets
used every day may outperform a precise one that is abandoned in a fortnight. Nothing here measures that,
and it may be the largest effect of all.
Frequently Asked Questions
How accurate are AI calorie counting apps?
Against doubly labelled water, the gold standard, a food-recognition app showed a bias of −329.6
kcal/day with limits of agreement from −1,503.8 to +844.5 kcal/day in 30 adult women over seven days.
Across the wider literature, a systematic review of 52 papers found relative error for calories ranging
from 0.10% to 38.3%. The honest answer is that accuracy varies enormously with the food and the
conditions, and no published study supports the ±1% figures that circulate in marketing.
What is doubly labelled water and why does it matter?
It measures energy through isotope tracking in the body during ordinary daily life, and it is the
reference method against which dietary assessment tools are validated. It matters because most app
accuracy claims compare the app against a nutrition database, which only tells you whether the app
matched a table. DLW compares against the person, which is the question you actually care about.
Is an app better than just guessing or writing it down?
In the study above, yes. The 24-hour dietary recall method used in nutrition research for decades
showed a larger bias of −543.0 kcal/day and wider limits of agreement, from −1,802.5 to +716.5. The
authors concluded the app gave a closer representation of energy intake than 24-hour recall. Every method
of measuring human food intake is imprecise; the app is currently the least imprecise of the practical
ones.
Why do apps do worse on some cuisines?
A 2024 study in Nutrients found manual logging apps overestimated energy for Western diets by an
average of 1,040 kJ and underestimated Asian diets by an average of −1,520 kJ. Mixed dishes, shared
plates, layered sauces and foods poorly represented in Western-built food databases are the recurring
problem. The error is directional rather than random, which means it is a database and training issue
rather than bad luck.
Should I trust benchmark sites that rank these apps?
Check three things before you do. Whether the numbers are internally consistent, because the same app
appears with different figures across pages in this space. Whether different benchmarks crown different
winners, which suggests preference rather than method. And whether the site’s claims about itself hold up:
one such benchmark describes itself as established in 1999, two decades before the technology existed. A
precision claim tighter than anything in peer review needs to carry its own proof.
Which app was most accurate in published research?
In the Nutrients assessment of 18 apps drawn from 800 screened, MyFitnessPal scored highest at 97%
accuracy and Fastic second at 92%, measured against food records for three diet types. Note that this
compares against food records rather than doubly labelled water, so it is answering a narrower question
than the DLW study does. It is also a 2024 snapshot of a field where apps change constantly.
Can I use one to hit an exact calorie deficit?
The published spread makes that unreliable. Limits of agreement of −1,504 to +845 kcal/day mean an
individual day’s figure carries limited information about that day. What survives the imprecision is the
trend: if the tool is consistently biased in one direction, the change over weeks is still readable even
when the daily number is not. Use it as a direction indicator rather than as a measurement.
Does any of this mean I should not use one?
No. Self-monitoring is among the better-supported behavioural components of weight management, and
consistency of use is likely to matter more than decimal-place accuracy. The purpose of knowing these
figures is not to abandon the tool but to stop treating its output as a measurement when it is an
estimate with a wide and documented range. This article summarises published research and is not medical
advice.
Sources
- Assessing daily energy intake in adult women: validity of a food-recognition mobile application
compared to doubly labelled water. Frontiers in Nutrition, 2023. PMID 37810925. 30 adult women
with normal body weight, seven days free-living; app 1,905.5 ± 531.1 kcal/day against DLW 2,235.2 ± 456.5
kcal/day; bias −329.6 kcal/day; limits of agreement −1,503.8 to +844.5 kcal/day; R² = 27%, p = 0.50;
24-hour recall bias −543.0 kcal/day with limits of agreement −1,802.5 to +716.5 —
Full text - AI-based digital image dietary assessment methods compared to humans and ground truth: a systematic
review. Annals of Medicine, 2023. PMID 38060823. 52 papers 2010-2023; convolutional neural
networks in 79%; relative error 0.10% to 38.3% for calories and 0.09% to 33% for volume; ground truth
from nutrient tables in 51% and weighed food in 27%; conclusion that tools need more development before
deployment as stand-alone dietary assessment methods —
PubMed record - Evaluating the Quality and Comparative Validity of Manual Food Logging and Artificial
Intelligence-Enabled Food Image Recognition in Apps for Nutrition Care. Nutrients, 2024.
PMID 39125452. 18 apps assessed from 800 screened; MyFitnessPal 97% and Fastic 92% accuracy; manual
logging overestimated Western diets by a mean 1,040 kJ and underestimated Asian diets by a mean −1,520 kJ
—
PubMed record - Comparative accuracy of smartphone apps and a generative AI tool for carbohydrate counting: an
independent bicentric study. PMID 41424217. Mean absolute errors in the region of 13 to 20 g per meal
across the tools tested —
PubMed record
All figures above are taken from the published abstracts of the
studies cited, checked August 2026. Where a claim circulating online could not be traced to a
peer-reviewed source, it is described as a marketing claim rather than reported as a finding.


