From Aggregation to Integration: Why Broad Data Beats Big Data
.png)
Agencies have accumulated more data than ever before, but aren't seeing the benefits from it. A typical agency might have five years of ATSPM data sitting in one silo, several third party data sources each reporting into a portal of their own, and crash data held somewhere else again, with every new product arriving under its own dashboard and its own login. All of that data has been paid for and is technically available, yet answering a single question still requires exporting from each system separately and stitching the pieces together manually. The result is a set of complicated, convoluted workflows that rarely reach an answer to a decision, so the hardest questions go unanswered. Whenever data is left on the table, money is left on the table with it.
This is a structural problem. Agencies have concentrated on aggregating their data for years, when the value comes from integrating it. The two are easy to confuse, so let's define them.
The difference between aggregation and integration
- Aggregation stores different datasets in the same place, so they can be retrieved from a single location.
- Integration goes a step further, giving those datasets a common data model so they can be processed together.
Aggregation is a genuine improvement over scattered systems, because it at least brings the data into one place, but since each dataset keeps its own structure, combining any two of them remains manual work that has to be repeated time after time. Integration removes that repeated effort, because once the data shares a common model it can be retrieved in a single step and presented in one interface, allowing a team to complete a task in full from one place instead of moving between systems and re-keying as they go.
The ability to process data together determines whether an agency can answer its questions at all, and that comes down to how transportation decisions are made.
Every real decision needs a combination of data
Every decision in transportation is complex, drawing on several dimensions at once, and no single dataset holds all the answers on its own. A sound decision almost always depends on a combination of data, and the value comes from assembling that combination easily and correctly. Aggregation can line those datasets up side by side, but only integration allows an engineer to use them together at the moment a decision is made.
It is also why collecting more of the same data has stopped paying off.
Broad data beats big data
When agencies cannot fully use the data they have, the usual response is to gather even more of it. In signal management, we regularly see agencies holding 5+ years of ATSPM data and still struggling to draw any value from it, because more of the same data does not change the picture. A deeper archive of a single source only answers the same narrow questions it always did. The picture changes when an agency adds different kinds of data and integrates them so they can be read together, which is why the diversity of the data matters far more than its volume.
The clearest way to see this is in one everyday decision, such as left-turn phasing.
Left-turn phasing: one decision, five datasets
Left turns are among the most dangerous movements at an intersection. In a NHTSA study of intersection-related crashes, a left turn was the critical event in 22.2 percent of them, compared with 1.2 percent for right turns, and the head-on and angle collisions a left turn produces are often among the most severe. That is why deciding how to phase a left turn is a call every agency has to get right: whether to run it protected, permissive, or protected-permissive, and when to change it. The answer relies on several factors.
ATSPM can tell you the current phasing type and where the gaps are, but it cannot tell you much beyond that, and to make the call you need several other inputs:
- Opposing movement speed, from segment-based probe data
- Control delay, from signal-based probe data
- Crash history, from GIS data
- Number of lanes, from mapping data
Those are five inputs drawn from five different datasets, and all of them feed a single decision. If the data is only aggregated, it sits in the same place but still has to be reconciled by hand, whereas if it is integrated, the whole picture appears in a single view and the decision can be made on all five inputs at once. This is why broad data beats big data every time.
Bigger isn't better, broader is
Collecting more of the same data will not move an agency forward, but integrating the data it already holds will. Integration is what turns a collection of separate feeds into better decisions.
This is the approach Flow Labs has pioneered over the past five years. Agencies working from integrated data now make calls like this one in minutes, not months. They get the full value of the data they already pay for, and they get more done.
If you would like to see it in action on your own network contact our team.
Manage every traffic management system and data in one platform, without the need for new hardware.

.avif)