25 years of improving predictions

NSF NCAR’s Data Assimilation Research Testbed continues to innovate

Aug 31, 2026 - by Laura Snider

For the last quarter century, the U.S. National Science Foundation National Center for Atmospheric Research (NSF NCAR) has been empowering the Earth system science community to make better predictions — about the weather and far beyond — by making the critical step of data assimilation easier for researchers, students, agencies, and others to implement.

The Data Assimilation Research Testbed (DART) is preparing to celebrate its 25th birthday in January. Over its lifetime, DART has been continuously improved and expanded, with more than 100 code releases. It now supports more than 50 models that simulate everything from high-resolution, local weather to the global state of the ocean. DART can be used with models designed to forecast streamflow, seasonal temperature, sea ice extent, air quality, and much more. In total, DART has contributed to research outcomes described in about 500 peer-reviewed publications and 50 student dissertations and theses. 

“DART is a true community facility,” said NSF NCAR scientist Moha Gharamti, one of DART’s lead developers. “Anyone can make a prediction, but not every prediction is good. Good data assimilation techniques are an important component of any good prediction system, and for the last 25 years, we have focused on making a world-class tool that lowers the barriers for researchers in our community working on predictive models.”

DART’s many advancements since the platform was first developed are outlined in a paper published last year in the Bulletin of the American Meteorological Society.  These accomplishments will be celebrated during a special session of the American Meteorological Society Annual Meeting, which will be held in Denver in January.

Estimating the state of the atmosphere

Even a perfect weather model could not accurately predict tomorrow’s weather if it doesn’t know the state of the atmosphere today, and the state of the atmosphere today is not fully known. Of course we measure the temperature, winds, humidity, and other variables at a number of ground-based weather stations, but that only gives us information about a single geographic point, and even then, the instruments used to make those measurements can have errors. 

Vast areas of the Earth do not have weather stations or other in-situ instruments that measure conditions where they are. These sparsely observed regions include the oceans, the poles, rural areas, and sizable swaths of the global South. 

Weather models require starting conditions at each point on a 3D grid, spreading horizontally over the globe or regional forecast area and vertically up into the atmosphere whether there are observations taken nearby or not. Data assimilation is the technique scientists use to estimate these values — a best guess at the current state of the atmosphere. 

Data assimilation combines models with observations, which include not only ground-based weather stations, but also data from satellites, weather balloons, dropsondes, radar, aircraft, ships, buoys, and more. Each of these observing platforms send data at different time intervals and each has different margins of error. Combining them to create a coherent and consistent state of the atmosphere is an exceptional challenge. 

Data assimilation also improves predictions by continuing to incorporate new observations as a model runs, nudging the model forecast back toward a more probable forecast as information comes in.

An algorithm breakthrough leads to a community facility

Data assimilation was developed in tandem with the first numerical weather prediction models, which were made possible by the birth of computers in the 1940s and ‘50s. Early data assimilation had to be programmed into the code of individual models, a laborious process.

However, advancements in the mathematics underpinning data assimilation in the 1980s and ‘90s opened the door for a more flexible approach. In 2001, NSF NCAR scientist Jeffrey Anderson improved further on current data assimilation techniques by introducing the Ensemble Adjustment Kalman Filter. This algorithm uses multiple forecast runs over the same time period, creating a spread of possible futures. Then it compares this spread with the observational data as it arrives, choosing a best guess for the state of the atmosphere that combines these sources of information. 

The efficiency and effectiveness of this new technique made it feasible to begin building a new, community-based platform that would make data assimilation much more widely accessible. The first version of DART was publicly released in April 2004. 

“In 2001, NSF NCAR had lots of great Earth system models but most of them had no data assimilation capability for a good reason: state-of-the-art variational data assimilation systems required person-decades of development for each model,” Anderson said. “We set out to develop an ensemble data assimilation framework that would require orders of magnitude less development time for each model while rivaling the best existing systems and providing additional information for uncertainty quantification. It seemed like a crazy idea at the time, providing assimilation for all of NSF NCAR’s flagship models, but it worked.”

One of DART’s great strengths is that it does not require any modification of the model’s code to run. If a researcher has a prediction model and a set of observations, it’s possible to use DART to combine those things. In fact it does not even matter what the model is predicting. While DART was developed for use by the Earth system science community, its possible applications are much broader. For example, during the pandemic, a DART interface was developed for an infectious disease model that simulated the progression of infections and the impact of vaccinations.

DART also has uses beyond enhancing predictions. Data assimilation can help assess what the impact of new satellites or other observational platforms would be before they are built, helping to direct investments. It can also help plan flight paths for aircraft releasing dropsondes — instrument packages that fall through storms — to ensure that the data that come back are as useful as possible. And data assimilation can be used to compare models and assess their relative strengths and weaknesses.

A platform for all

DART’s algorithms have been updated over the decades, and recent progress has allowed the system to be used with ultra high-resolution simulations and with artificial intelligence models. DART developers have also continued to improve their training, tutorials and documentation. The DART team is focused on making the tool as easy to learn and use as possible, enabling all kinds of high-impact research.

“We actively invite new collaborators who bring new models, new observations, and new applications to us,” Gharamti said. “Whether you are an academic or a government researcher, operational forecaster, or a student exploring data assimilation for the first time, we welcome your ideas and contributions.”

About the article

Title: The Data Assimilation Research Testbed: A Robust, Scalable Software Facility with Groundbreaking Capabilities for Model-Data Integration
Authors: Mohamad El Gharamti, Helen Kershaw, Kevin Raeder, Brett Raczka, Benjamin Johnson, Marlena Smith, Jeffrey Anderson, Daniel Amrhein, Nancy Collins, Timothy Hoar, Benjamin Gaubert, Ian Grooms, and Lukas Kugler
Journal: Bulletin of the American Meteorological Society

See all News