Time-Series Foundation Models for Industrial Sensors: TimesFM vs Chronos vs Moirai
A pump bearing starts to degrade on a Tuesday, and your historian has three years of its temperature and vibration features. The pitch for time series foundation models is that you no longer need to train a bespoke forecaster on that history: you hand a pretrained model the last few thousand samples, and it returns a calibrated forecast with no training step. For retail demand and electricity load that pitch has largely held up in public benchmarks. For a machine that spends 99 percent of its life in steady state and produces its most interesting data in the last 1 percent, it is far less obvious.
The field also moved quickly in the last twelve months. Amazon shipped Chronos-2 with covariate support, Salesforce released a much smaller Moirai 2.0, and Google released TimesFM 3.0 on 31 August 2026 with native multivariate attention but a non-commercial weights licence. That licence change alone should alter how an industrial team shortlists models.
This post compares TimesFM, Chronos and Moirai at the mechanism level: how each turns a sensor window into a forecast, what the current versions and licences actually are, where zero-shot forecasting breaks on vibration and temperature data, and how to run an evaluation that tells you whether to ship. You leave with a decision matrix and a protocol you can run on your own assets.
What this covers: the architectures and current versions of the three families, licence and deployment constraints, a zero-shot evaluation protocol for sensor data, honest failure modes, and a recommendation checklist for predictive maintenance teams.
Context and Background
Until about 2023, forecasting in industrial settings meant one of two things. You fit a statistical model per signal (exponential smoothing, ARIMA, a seasonal naive baseline), or you trained a neural network per asset class on your own history. Both approaches share a cold-start cost: every new sensor, line or plant needs enough history and enough engineering time before the forecast is trustworthy. Anyone who has deployed condition monitoring across a fleet knows that cost scales with the number of distinct machines, not the number of data points.
Foundation models attack that cost the way language models did for text. You pretrain once on a huge and varied corpus of series, then apply the model to unseen series with no gradient updates. Three research lines defined the first generation. Google’s TimesFM paper (Das, Kong, Sen and Zhou, arXiv 2310.10688) described a patched decoder-only transformer. Amazon’s Chronos (arXiv 2403.07815) tokenized values into a fixed vocabulary and reused the T5 language-model family. Salesforce’s Moirai (arXiv 2402.02592) trained a masked encoder on LOTSA, an archive the authors describe as more than 27 billion observations across nine domains.
Each line has since been revised. As of the sources I checked for this post, the current picture is: TimesFM 3.0 (August 2026, weights under a non-commercial licence), TimesFM 2.5 (200M parameters, Apache 2.0), Chronos-2 (October 2025, 120M parameters, Apache 2.0), Chronos-Bolt (November 2024, 9M to 205M parameters), and Moirai 2.0 (small variant, released August 2025, 11.4M parameters). The Moirai licence is inconsistent across Salesforce’s own pages, which I cover below.
Benchmarks have kept pace. GIFT-Eval from Salesforce and fev-bench from the AutoGluon team are the two leaderboards that most new model announcements cite. The fev-bench paper describes 100 forecasting tasks across seven domains, 46 of them with covariates, and reports win rates and skill scores with bootstrapped confidence intervals. These are useful for ranking general-purpose models. They are not a substitute for testing on your machines, and I will argue later that for industrial signals the gap between leaderboard and plant floor is the central risk.
If you are working on the data side first, the earlier posts on time-series database platform architecture for industrial IoT and on running time-series forecasting at the edge in production cover storage, windowing and deployment patterns that this comparison assumes. For the maintenance use case itself, see the comparative analysis of machine learning models for predictive maintenance. For the primary sources on the three model lines, start with the TimesFM repository, the Chronos forecasting repository and the Salesforce uni2ts repository.
One framing point before the architecture. A forecasting model is not an anomaly detector, and not a predictive maintenance system. It predicts the next values of a signal. Anomaly detection arrives only when you compare that prediction with reality and decide the gap is meaningful, and predictive maintenance arrives only when you connect that gap to a failure mode and a lead time. Foundation models improve the first step. The other two steps are still your engineering.
Core Architecture: How Each Family Turns Sensor Windows into Forecasts
The three families differ in how they represent a numeric window to a transformer: Chronos quantizes values into discrete tokens and treats forecasting as language modelling, TimesFM groups values into patches and predicts patches with a decoder-only stack, and Moirai pairs patches with a variate-aware attention scheme and, since version 2.0, a decoder-only design with multi-token quantile output. These choices drive latency, memory, and how each model fails.

Figure 1: Three input paths into a forecast. All three end in a quantile or sample forecast whose residual against the real signal becomes the anomaly score.
Figure 1 shows the shared skeleton. Every model scales the window, converts it into a sequence of tokens or patches, runs a transformer, and emits a distribution over future values. The residual between that distribution and what the sensor then reports is the raw material for anomaly scoring. The differences are in the middle two boxes.
Chronos: values as a vocabulary
The original Chronos (March 2024) scales each series and quantizes the scaled values into a fixed vocabulary of bins, then trains T5-family models with a cross-entropy loss to predict the next token. The paper describes models from 20M to 710M parameters, trained on a mix of public datasets and a synthetic corpus generated from Gaussian processes (the KernelSynth procedure). The Hugging Face collection lists the family as Tiny (8M), Mini (20M), Small (46M), Base (200M) and Large (710M), with a 512-step context. Sampling several trajectories autoregressively yields the forecast distribution.
The consequence for industrial data is mostly about resolution. Quantization into a finite number of bins bounds the precision with which a value can be represented after scaling, and scaling is computed from the context window. If a window contains a single large spike, the scale inflates, the bins coarsen, and the fine structure of normal operation shrinks into a few adjacent tokens. That is a design property, not a bug, but it is exactly the situation that a bearing impact produces in a vibration channel.
Chronos-Bolt (November 2024) changed the mechanism. It uses patches of the input as tokens and decodes several future steps directly rather than autoregressively. The repository describes four sizes, Tiny (9M), Mini (21M), Small (48M) and Base (205M), and claims up to 250 times faster inference and 20 times better memory efficiency than the original Chronos models of the same size. That is a vendor claim measured by the authors against their earlier family; treat it as directional, and measure on your own hardware.
Chronos-2 (October 2025) is a different architecture again. According to its model card it is a 120M-parameter encoder-only model inspired by the T5 encoder, with a maximum context of 8,192 and a maximum prediction horizon of 1,024 steps. Its distinguishing feature is a group attention mechanism that lets the model attend across a set of related series and covariates inside one forward pass. It handles univariate, multivariate and covariate-informed tasks, with past-only and known-future covariates (both real-valued and categorical). The card reports more than 300 forecasts per second on a single A10G GPU and supports CPU inference. The licence is Apache 2.0.
TimesFM: patches and a decoder
TimesFM treats a window as a sequence of fixed-length patches. An input patch is embedded by a small residual network, passed through stacked causal transformer layers, and the output at each position predicts the next, longer patch. Predicting a longer output patch than input patch is the key efficiency trick in the paper: a horizon of 128 steps needs far fewer autoregressive iterations than one step at a time.
The version history matters here, because the public numbers differ by release. TimesFM 2.0 was a 500M-parameter model with a 2,048-step context. TimesFM 2.5, released on 15 September 2025, cut the model to 200M parameters, extended the context to 16k according to the repository README, removed the frequency indicator that earlier versions required, and added an optional 30M-parameter quantile head for continuous quantile forecasts. Covariate support through an XReg (external regressor) path was added in October 2025. One caveat: the Hugging Face model card for the 2.5 PyTorch checkpoint lists a maximum context of 1,024 and a maximum horizon of 256 in its configuration notes, while the README says 16k context and 1k horizon. I could not reconcile the two from the documents; check the max_context value in the version you install rather than trusting either page.
TimesFM 3.0, announced on 31 August 2026, is the largest change. Its model card describes a 0.3B-parameter “Stacked Mixing Transformer with Variate Attention and CPM Iterative RevIN” with 20 layers, a model dimension of 1,280, 16 attention heads, a context patch length of 32 and a forecast patch length of 64. Press coverage describes alternating temporal (causal, within-series) and cross-series attention layers, and a trillion-plus time points of pretraining data, though the training volume comes from reporting rather than a primary document I could open. Pretraining sources named on the card are GiftEvalPretrain (excluding overlaps with fev-bench), Wikipedia Pageviews through November 2023, Google Trends top queries through the end of 2022, and synthetic and augmented data. The repository lists top rankings on fev-bench, the TIME benchmark and GIFT-Eval, but the pages I read give no numeric scores, so I report none.
The licence is the practical headline. The source code is Apache 2.0, but TimesFM 3.0 weights are distributed under a separate timesfm-non-commercial-license-v1.0, which the repository describes as restricted to non-commercial, non-production use. If you want a model inside a production maintenance workflow, TimesFM 3.0 is currently a research and evaluation artefact, and TimesFM 2.5 remains the Apache-licensed option in that family.
Moirai: any-variate attention, then a rethink
Moirai 1.0 (arXiv 2402.02592, ICML 2024 oral) tackled three problems the authors name: learning across frequencies, handling a variable number of variates, and fitting heterogeneous distributions. It used multiple patch sizes tied to frequency, an any-variate attention mechanism that flattens variates into one sequence with learned identity biases, and a mixture-of-distributions output head. Moirai 1.1 followed in June 2024, and a sparse Mixture-of-Experts variant, Moirai-MoE, in October 2024.
Moirai 2.0 reverses several of those choices. The paper, “Moirai 2.0: When Less Is More for Time Series Forecasting” (arXiv 2511.11698), describes a decoder-only model trained on 36 million series that swaps masked-encoder training and mixture-distribution outputs for quantile forecasting with multi-token prediction. The authors claim it is 30 times smaller than Moirai 1.0-Large and twice as fast while performing better, and that it ranks among the top pretrained models on GIFT-Eval. They also state that performance plateaus with increasing parameter count and declines at longer horizons. That last admission matters for maintenance, where the horizon you care about is often long relative to the context.
The released small variant has 11.4M parameters. Its Hugging Face card describes patch embeddings that carry missing-value information, patch-level random masking for robustness, and a data-filtering step that removes series judged non-forecastable. On licensing, I found a conflict: the uni2ts GitHub README lists Apache 2.0 for its releases, while the Hugging Face cards for Moirai 2.0-R-small and Moirai 1.1-R-large state CC-BY-NC-4.0, and the 1.1 card describes the model as for research purposes in support of an academic paper. Until Salesforce reconciles the two, treat Moirai weights as non-commercial and ask for written confirmation before any production use.
What the architectures imply for sensor data
Three architectural facts matter more than parameter counts. First, context length: a 512-step window of one-minute features covers under nine hours, whereas 8,192 steps covers about five and a half days. Second, whether the model accepts covariates, because for an industrial asset the load, speed or ambient temperature often explains most of a signal. Third, whether the output is a calibrated distribution, because anomaly thresholds built on a quantile band are easier to defend than ones built on a single point forecast.
By that measure, Chronos-2 and TimesFM 3.0 are the families with native cross-channel mechanisms, TimesFM 2.5 supports covariates through its XReg path, and Moirai 2.0 is the smallest and cheapest of the group. The comparison table in the next section collects the numbers.
Deeper Analysis: Putting Zero-Shot Forecasters on Vibration and Temperature Data
The facts above come from model cards and papers. What follows is engineering reasoning about how those mechanisms meet industrial signals. Where I give numbers, they are labelled illustrative arithmetic, not measurements.
Version and licence comparison
| Model | Params | Context (documented) | Covariates | Output | Licence of weights |
|---|---|---|---|---|---|
| Chronos (2024) | 8M to 710M | 512 | No | Sampled token paths | Apache 2.0 |
| Chronos-Bolt | 9M to 205M | 2,048 | No | Direct quantiles | Apache 2.0 |
| Chronos-2 | 120M | 8,192 | Past and known-future, real and categorical | Quantiles, multivariate | Apache 2.0 |
| TimesFM 2.0 | 500M | 2,048 | No | Point forecast, not verified | Apache 2.0 |
| TimesFM 2.5 | 200M (+30M optional quantile head) | 16k per README, 1,024 on card | XReg path (Oct 2025) | Mean and 10th to 90th quantiles | Apache 2.0 |
| TimesFM 3.0 | about 0.3B | Not stated on pages read | Native multivariate and covariates | Quantiles | Non-commercial |
| Moirai 2.0 small | 11.4M | Not stated on card | Not stated on card | Quantiles, multi-token | CC-BY-NC-4.0 on card, Apache-2.0 on GitHub |
Context figures for Chronos-Bolt and Chronos-2 come from the Chronos-2 model card’s comparison (2,048 and 8,192). Empty or hedged cells mean I could not verify the value from a primary source.
Why sensor data differs from the benchmark data
Public pretraining corpora are dominated by business and environmental series: retail sales, web traffic, energy load, weather, transport counts, with daily, hourly and sub-hourly sampling. They are rich in trend and calendar seasonality. Industrial sensor data has a different statistical personality, and four differences do most of the damage.
First, regime structure. A compressor cycles between loaded, unloaded and stopped. The signal is a mixture of piecewise-stationary regimes with abrupt switches that follow a control program, not a calendar. A foundation model sees the switch as a level shift and has no way to know that it was commanded. Without a covariate such as an operating-state tag, the forecast around a switch is poor, and the residual looks like an anomaly even though nothing is wrong.
Second, sampling rate and bandwidth. Raw vibration is sampled in kilohertz, and the diagnostic information lives in spectral structure: bearing defect frequencies, gear mesh harmonics and sidebands. None of these models is designed to forecast a 25 kHz waveform sample by sample, and none of the pretraining corpora described above contains much of that kind of data. The practical route is to compute condition indicators first (RMS, kurtosis, crest factor, band energies, envelope-spectrum amplitudes) at a rate of seconds to minutes, then forecast those slowly varying features. That moves the problem into the territory these models were trained on.
Third, rarity and label scarcity. Failures are scarce. A forecasting model trained on a corpus where outliers are noise, and where Moirai 2.0 explicitly filters out series judged non-forecastable, has no particular reason to model failure onset. The positive class you actually care about is almost absent from both pretraining and your own history.
Fourth, quantization of the signal at the sensor. Many temperature and pressure channels are reported at coarse resolution with deadbands, so the series is a staircase with long flat stretches. A flat window is trivial to extrapolate, and any model will do well on a standard error metric. The metric then rewards predicting nothing, which is the next problem.

Figure 2: A pipeline that respects the models’ training distribution. Raw vibration becomes slow condition indicators, aligned with load and temperature, gated for data quality, forecast with quantile bands, and then scored with a persistence rule.
Figure 2 is the architecture I would put in front of any of the three models. The quality gate in the middle is not decoration. Foundation models are trained to be robust to noise, which also means they will happily forecast a flatlined sensor or a window full of interpolated gaps. Detect those conditions first and route them to a data-quality alert, not an anomaly alert.
A worked illustration of the horizon problem (illustrative arithmetic)
Suppose a bearing temperature feature is sampled once a minute and the degradation trend you care about develops over about 36 hours, which is 2,160 steps. A model with a 512-step context sees under nine hours of history, so the slow drift looks nearly flat inside the window. A model with an 8,192-step context sees about 5.7 days and can place the drift against several load cycles.
Now consider the output side. If you want a 6-hour-ahead forecast at one-minute resolution, that is 360 steps. Models that predict long output patches, such as TimesFM with a 64-step forecast patch in version 3.0, need only a handful of decoding passes; autoregressive token models need many more. The Moirai 2.0 authors report that accuracy declines at longer horizons, so a 360-step forecast is where the three families are most likely to diverge on your data. The pragmatic response is to downsample to five-minute features, which turns 360 steps into 72 and extends the same context window to roughly 28 days for 8,192 points. These numbers are arithmetic on assumed sampling rates, not benchmark results.
Anomaly detection by forecast residual
The standard recipe for turning a forecaster into an anomaly detector has three parts. Predict a quantile band for the next window. Compute a residual, either the signed distance outside the band or a normalised error such as the absolute error divided by the in-sample seasonal naive error (the MASE denominator). Then apply a persistence rule so that a single breach does not page anyone.
Quantile output is what makes this work. A point forecast gives you only a residual to threshold, and you must pick a threshold per signal. A 10th to 90th quantile band gives a natural, if crude, per-window uncertainty estimate. TimesFM 2.5 returns the mean and the 10th to 90th quantiles, Chronos-2 and Moirai 2.0 return quantiles directly, and the original Chronos requires sampling many paths and computing empirical quantiles at higher cost.
The caveat is calibration. A nominal 10 to 90 band should contain about 80 percent of observations in normal operation. On plant data it routinely does not, because the pretraining mixture does not match your regime structure. Before trusting any band, measure coverage on a held-out normal period and adjust, for example by conformal calibration on that period. If the band covers 95 percent when it should cover 80, your detector is deaf; if it covers 60 percent, you will drown in false alarms.
An evaluation protocol that tells you whether to ship
Leaderboard rank answers the wrong question. The right question is whether a model, given your signals, beats a cheap baseline by enough to justify its operational cost. Figure 3 outlines the protocol I recommend.

Figure 3: Evaluation protocol. Time-based splits feed baselines, zero-shot models and a light fine-tune; forecast metrics are necessary but the decision gate uses event metrics, lead time and false alarms.
The steps, in order:
- Split by time, per asset. Train-validation-test boundaries must be chronological. Random window splits leak adjacent samples and flatter every model.
- Always include baselines. Seasonal naive, exponential smoothing and a simple ARIMA or AutoETS fit. If a foundation model cannot beat seasonal naive by a clear margin on MASE, stop.
- Run zero-shot first, then add covariates. Give the same window to each candidate with and without operating-state and load tags. The gain from covariates tells you how much of your signal is controlled, not random.
- Measure forecast quality with scale-free metrics. MASE for point accuracy, and a quantile loss such as weighted quantile loss for distribution quality, with coverage reported separately.
- Replay faults as events. For each historical failure, record the earliest time the detector raised a persistent alert before the failure, and count the alerts per asset per month during healthy periods. Lead time and false alarms per month are the metrics that map to maintenance cost.
- Price the decision. Convert false alarms into inspection hours and lead time into avoided downtime, using your own numbers. This is where a small model with a worse benchmark rank often wins.
- Check stability. Rerun on a different month and a different asset of the same class. A model whose ranking flips between months is not ready.
Two practical details. Report results with confidence intervals; fev-bench’s use of bootstrapped intervals is the right model for this. And report per-asset results, not only a fleet average, because a foundation model that is excellent on 80 percent of assets and useless on the 20 percent with irregular duty cycles will look fine in aggregate and fail in production.
Decision matrix
| Scenario | First pick | Why | Watch out for |
|---|---|---|---|
| Fleet of pumps, one-minute features, load tag available | Chronos-2 | Covariates and 8,192 context, Apache 2.0 | GPU cost for large fleets |
| CPU-only edge gateway, univariate temperature trend | Chronos-Bolt Small or TimesFM 2.5 | Small, fast, permissive licences | No native cross-channel reasoning in Bolt |
| Research benchmark of the best available zero-shot accuracy | TimesFM 3.0 | Reported top ranks on fev-bench, TIME, GIFT-Eval | Non-commercial weights only |
| Tiny memory budget, many series, short horizons | Moirai 2.0 small | 11.4M parameters | Licence ambiguity, horizon decay |
| Long-horizon degradation trend with sparse failures | None zero-shot alone | Use forecast residual as a feature in a supervised or survival model | Few labelled failures |
The last row is deliberate. For prognosis, remaining-useful-life estimation needs a degradation model linked to failure labels. A forecaster supplies features; it does not supply the failure threshold.

Figure 4: Selection flow. Licence is the first gate, covariate need the second, and compute budget the third; every path ends in validation on your own data.
Figure 4 encodes the same logic as the matrix. Notice that licence comes first. A model you cannot legally deploy has no accuracy advantage.
Trade-offs, Gotchas, and What Goes Wrong
The metric rewards doing nothing. On a staircase temperature signal, every model scores well because the last value is a strong predictor. A leaderboard-style average error then hides the fact that the model contributes nothing at the moments that matter. Always report error on the windows around operating-state changes and around known fault onsets, not only over the whole history.
Level shifts are mistaken for faults. A setpoint change, a shift start or a product changeover produces a step that the model cannot predict without a covariate. Feed operating-state and setpoint tags as covariates where the model supports them (Chronos-2 natively, TimesFM through XReg or version 3.0’s covariate support), or suppress alerts for a fixed period after commanded transitions.
Scaling and normalisation interact with spikes. Every family normalises the context window, in some form, before the transformer sees it. A single outlier in the context can change the scale for the whole window. For Chronos-style quantization this also coarsens the bins. Winsorise or clip obvious sensor glitches before inference, and keep the raw value for the alert record.
Silent distribution shift after maintenance. After a bearing replacement, the new normal differs from the old normal. A zero-shot model adapts through its context window, which is an advantage over a trained-per-asset model, but the context contains the old regime until it ages out. Reset or truncate the context at logged maintenance events.
Missing data and irregular sampling. Historians record on change, with timestamps that are not on a grid. These models expect regular sampling. Resample to a fixed interval with an explicit missing-value policy, and be aware that only some models represent missing values explicitly (the Moirai 2.0 card describes patch embeddings that carry missing-value information). Forward-filling a long gap creates a flat segment that looks like a stuck sensor.
Licence risk is a real engineering constraint. Some of the most prominent current checkpoints, namely TimesFM 3.0 and the Moirai checkpoints as labelled on their Hugging Face cards, carry non-commercial terms. Building a production pipeline on one and discovering this at procurement time is an expensive mistake. Record the licence of every model weight you ship in your software bill of materials.
Benchmark contamination and leakage. Pretraining corpora overlap with public benchmarks to varying degrees; the TimesFM 3.0 card says it excludes overlaps with fev-bench from GiftEvalPretrain, which is a good sign, but you should still assume that popular public datasets may have appeared in pretraining of any model. Your own plant data is the clean test set, which is a reason to hold it back and not to tune on it casually.
Compute and ops cost. A 120M to 330M model is small by language-model standards but not trivial at the edge. The Chronos-2 card reports over 300 forecasts per second on an A10G, which is ample for a fleet scoring every minute, but a CPU-only gateway scoring thousands of channels needs either the smaller Bolt or Moirai checkpoints, or batching. Measure end-to-end latency including preprocessing, since feature extraction often costs more than inference.
Anomaly detection is not forecasting. Several recent surveys, including a 2024 review of foundation models for anomaly detection and prediction (arXiv 2412.19286), note that forecasting-based detection inherits the weaknesses of the underlying forecaster. A model that forecasts well can still miss a slowly developing fault, because a slow drift is easy to forecast by extrapolating it. Drift-type faults are better caught by comparing against a population of peer machines than by forecasting one machine against itself.
Fine-tuning is possible but changes the economics. The TimesFM repository added a LoRA fine-tuning example with Hugging Face Transformers and PEFT in April 2026, and Chronos models can be fine-tuned. Fine-tuning on a handful of failure windows risks memorising them. If you need it, validate on assets held out entirely from the tuning set.
Practical Recommendations
Start with the licence and the data, not the leaderboard. If production use is required, shortlist the Apache-licensed checkpoints: Chronos-2 for covariate-heavy fleets, Chronos-Bolt for CPU-constrained gateways, and TimesFM 2.5 as a second opinion. Use TimesFM 3.0 and Moirai for research comparison until their terms allow production use, and confirm Moirai’s position in writing.
Forecast condition indicators, not raw waveforms. Compute RMS, kurtosis, crest factor and band energies at the edge, align them with load and speed tags, and forecast those at one- to five-minute resolution. Build a quality gate that diverts flatlines, gaps and sensor glitches to a data-quality queue.
Treat the forecaster as one feature generator in a larger pipeline. Quantile residuals, band breach counts and coverage drift make good inputs to a supervised detector, a survival model or a rule engine that maintenance planners already understand. For model families beyond forecasting, such as classification of fault signatures, the machine learning model comparison for predictive maintenance and the condition monitoring architecture post cover that terrain.
Run the protocol in Figure 3 on three to five representative assets before any fleet rollout, and keep the baseline in production as a fallback. A seasonal naive forecast costs nothing and gives you a sanity check that triggers an alert if the foundation model’s error suddenly exceeds it.
Checklist
- [ ] Licence of each candidate weight file recorded and cleared for commercial production use.
- [ ] Raw waveforms reduced to condition indicators before forecasting.
- [ ] Regular-grid resampling with an explicit missing-value policy.
- [ ] Operating-state, load and setpoint tags supplied as covariates where supported.
- [ ] Chronological splits per asset; seasonal naive and ETS baselines included.
- [ ] Quantile coverage measured on a held-out normal period and recalibrated if off.
- [ ] Event metrics reported: lead time before failure and false alarms per asset per month.
- [ ] Context reset at logged maintenance events.
- [ ] Latency measured end to end on the target hardware.
- [ ] Fallback baseline kept running alongside the model.
Frequently Asked Questions
What are time series foundation models?
Time series foundation models are large neural networks, usually transformers, pretrained on very large and varied collections of time series so that they can forecast new series without per-series training. You supply a recent window of values and get back point or quantile forecasts. Examples include TimesFM from Google, Chronos from Amazon, and Moirai from Salesforce. Their main advantage is removing the cold-start training step for each new sensor; their main limitation is that pretraining data rarely resembles industrial signals such as vibration or regime-switching process variables.
Which is best for industrial sensor data: TimesFM, Chronos or Moirai?
No family wins everywhere, and the answer depends on licence first. Chronos-2 is Apache 2.0, accepts past and known-future covariates and has an 8,192-step context, which suits fleets with load and operating-state tags. TimesFM 2.5 is Apache 2.0 and compact, while TimesFM 3.0 reports top benchmark ranks but ships non-commercial weights. Moirai 2.0 is the smallest at 11.4M parameters but its licence is inconsistent across pages. Benchmark the permissively licensed candidates on your own assets before choosing.
Can foundation models detect anomalies in vibration data?
Not directly on raw waveforms. These models forecast slowly varying series, and raw vibration sampled in kilohertz carries its diagnostic information in spectral structure they were not trained to model. The workable approach is to compute condition indicators such as RMS, kurtosis and band energies, forecast those, and flag persistent residuals that fall outside a quantile band. Evaluate with lead time and false alarms per month, and confirm band calibration on a normal period before trusting any threshold.
What does zero-shot forecasting mean, and when does it fail?
Zero-shot means the model forecasts a series it never saw in training, using only the supplied context window and no gradient updates. It tends to fail on regime-switching machines where operating-state changes are commanded rather than seasonal, on rare failure onsets absent from pretraining data, on staircase signals where trivial baselines already score well, and when context windows are too short to show a slow degradation trend. Adding covariates and truncating context at maintenance events reduces some of these failures.
Are these models free to use commercially?
It depends on the checkpoint. Per the repositories and model cards I read, Chronos, Chronos-Bolt and Chronos-2 are Apache 2.0, and TimesFM 2.0 and 2.5 are Apache 2.0. TimesFM 3.0 weights use a non-commercial licence restricted to non-production use, even though the source code is Apache 2.0. Moirai checkpoints show CC-BY-NC-4.0 on their Hugging Face cards while the GitHub README lists Apache 2.0, so obtain written clarification from Salesforce. Licences can change between releases, so re-check at deployment time.
Do I need a GPU to run them?
Not necessarily. The Chronos-2 card says it supports both GPU and CPU inference and reports more than 300 forecasts per second on one A10G GPU. Smaller models such as Chronos-Bolt Tiny at 9M parameters and Moirai 2.0 small at 11.4M are practical on CPU-only gateways at modest series counts. Measure end-to-end latency on your target hardware, including preprocessing, because feature extraction and data movement often dominate the cost of the model itself.
Further Reading
- Time-series forecasting at the edge in production for deployment, windowing and monitoring patterns.
- A comparative analysis of machine learning models for predictive maintenance for the supervised and survival models that sit downstream of a forecaster.
- AI and ML for predictive maintenance: careers guide for the skills this work requires.
- AI weather forecasting models: GraphCast, GenCast and Aurora for another case of foundation-style forecasting on physical systems.
- Condition monitoring architecture for machinery health for the sensing and diagnostics layer.
- fev-bench: A Realistic Benchmark for Time Series Forecasting for a rigorous model of how to report forecasting results.
References
- Das, Kong, Sen, Zhou. “A decoder-only foundation model for time-series forecasting.” arXiv 2310.10688. https://arxiv.org/abs/2310.10688
- Amazon Science. “Chronos: Learning the Language of Time Series.” arXiv 2403.07815. https://arxiv.org/abs/2403.07815
- “Chronos-2: From Univariate to Universal Forecasting.” arXiv 2510.15821. https://arxiv.org/abs/2510.15821
- Woo et al. “Unified Training of Universal Time Series Forecasting Transformers.” arXiv 2402.02592. https://arxiv.org/abs/2402.02592
- Liu et al. “Moirai 2.0: When Less Is More for Time Series Forecasting.” arXiv 2511.11698. https://arxiv.org/abs/2511.11698
- “fev-bench: A Realistic Benchmark for Time Series Forecasting.” arXiv 2509.26468. https://arxiv.org/abs/2509.26468
- Google Research TimesFM repository and model cards (timesfm-2.5-200m-pytorch, timesfm-3.0-pytorch). https://github.com/google-research/timesfm
- Amazon Science Chronos forecasting repository and Chronos-2 model card. https://github.com/amazon-science/chronos-forecasting and https://huggingface.co/amazon/chronos-2
- Salesforce AI Research uni2ts repository and Moirai model cards. https://github.com/SalesforceAIResearch/uni2ts
- The New Stack, coverage of the TimesFM 3 launch, 31 August 2026. https://thenewstack.io/google-timesfm-3-multivariate-forecasting/
- “Time Series Foundational Models: Their Role in Anomaly Detection and Prediction.” arXiv 2412.19286. https://arxiv.org/pdf/2412.19286
By Riju – about
