HomeAnswers — NASA battery dataset
Technical question · Datasets

What can you actually prove with the NASA battery dataset?

That a method behaves as described. Not a failure rate, not a fleet statistic, and not a claim about modern cells. The NASA PCoE set is small, old, and full of capacity regeneration — which makes it an excellent instrument for checking methodology and a poor one for making performance claims.

The short answer

It is the most-used battery prognostics dataset in the field and the most over-claimed. The cells are 18650s aged more than a decade ago; modern commercial cells generally last longer and degrade differently. Rest periods produce capacity regeneration — recovery spikes that are physically real and that a model can be quietly tuned to smooth away. And the retained cell counts in most published work are small enough that a single cell moves the headline number several points. None of that makes it a bad dataset. It makes it a methodology instrument.


How we use it, and what we say about it

Six retained cells is an anecdote.

We run a blind hindcast: forecasts are made only from data available at that moment, and every cell is scored.

Including the one where the method never called a crossing at all, which we show in red. That cell is on our Evidence page and in the live demo, and it stays there.

The scoring is per checkpoint rather than per cell, so a model that gets one cell exactly right and three cells late cannot hide behind an average. Most checkpoints land inside our tolerance. Not all of them do.

We report lead time in cycles, never in days. Days require an assumed duty cycle, and an assumed duty cycle is how a cycle-domain result quietly becomes a calendar-domain promise.

And we say the size out loud: a handful of retained cells is an anecdote, not a fleet failure rate. Validation is moving to a larger commercial cohort for exactly that reason.

The three ways this dataset produces results that do not survive contact with a real pack.

Splitting by cycle instead of by cell. If cycles from the same cell appear in both training and test, the model has seen that cell’s idiosyncratic ageing signature. The number that comes out is optimistic and the leakage is easy to miss.

Smoothing away capacity regeneration. The recovery spikes after rest are a real physical phenomenon. Filtering them makes curves prettier and makes the model worse at the thing that matters, which is calling a threshold crossing near the end.

Generalising from a handful of devices. The companion IGBT accelerated-aging set has six devices. You cannot state a generalisable remaining-life claim from six devices, and neither can we.


Which dataset for which question

If you need something the NASA set cannot give you.

MIT/Stanford (Severson–Attia). 124 commercial LFP cells, 72 fast-charge protocols — but a single chamber temperature and one manufacturer and batch. Excellent for charge-protocol questions, no use at all for temperature extrapolation.

Protocol variation

Sandia (Preger et al.). Multiple chemistries across a range of chamber temperatures. This is the set to use when the question is whether a model holds outside the temperatures it trained on.

Temperature variation

ISU-ILCC. A designed multi-stage experiment, useful as an independent second source rather than a substitute.

Independent check

NASA PCoE. Small, old, regeneration-heavy. Ideal for demonstrating that a method does what you say it does, in public, where anyone can check it.

Methodology

A caution that applies across all of them: published work by Schauser and colleagues in Frontiers in Energy Research in 2022 found models transferring poorly across datasets, with a complex model producing unphysical lifetime predictions while a simpler variance-based model still erred by 17 to 30%. Cross-dataset generalisation in battery prognostics is not close to solved, and any vendor implying otherwise is worth questioning.

What this does not show
Our reference-cell work is on public data. It demonstrates methodology; it is not a performance guarantee.
Per-cell error figures and pass rates stay internal until the larger cohort lands, and are available under diligence.
Choir has no deployed customer reference and no field case study.

See the same discipline on your own assets.

A read-only assessment runs on data you already have, before any hardware conversation. You see the record; you decide what it is worth.

Early access · software-first · every number traces to a dated report