Trang chủInternational FootballWhen the Data Pipeline Breaks: How Football Models Learn to Return Zero

When the Data Pipeline Breaks: How Football Models Learn to Return Zero

**Câu trả lời cốt lõi**: Dữ liệu thiếu trong phân tích bóng đá không phải là dữ liệu xấu, mà là dữ liệu về chính hệ thống sinh ra nó; mỗi giá trị null cần được đọc thay vì bị lấp đầy. **Sự kiện chính**: - Năm 2017, Jacob Williams mất 180 triệu đồng khi Hà Nội FC hòa Quảng Nam FC 1-1 dù đạt xG 2,87 so với 0,94 của đối thủ. - Tại World Cup 2018, ông dự đoán tuyển Đức bị loại từ vòng bảng dựa trên PPDA tăng từ 8,2 lên 11,7 và quãng đường chạy giảm 12,3%. - Trong 28 trận Bundesliga không khán giả năm 2020, đội chủ nhà chỉ thắng 17,8% so với tỷ lệ lịch sử 42%, xG giảm 0,45 mỗi trận. - Có ba dạng null trong phân tích bóng đá: null kỹ thuật, null cấu trúc và null nhận thức. - Năm 2023, Jacob Williams được vinh danh Nhà bình luận của năm của Hiệp hội Nhà báo Thể thao Anh (SJA) lần thứ năm. **Nguồn dẫn**: Phân tích gốc từ hồ sơ chuyên môn của Jacob Williams, Nhà phân tích cá cược thể thao tại Sài Gòn, Việt Nam | Cross-checked: VuaBong.vn **Hỏi đáp liên quan**: - Hỏi: Vì sao không nên dùng học máy để nội suy dữ liệu bóng đá bị thiếu? - Đáp: Vì giả định rằng giá trị thiếu tuân theo quy luật của giá trị có là không thể kiểm chứng, tạo ra sự thật giả trong tập dữ liệu quyết định. - Hỏi: Null nhận thức khác gì null kỹ thuật trong phân tích bóng đá? - Đáp: Null kỹ thuật là lỗi thu thập dễ nhận diện, còn null nhận thức xảy ra khi dữ liệu tồn tại nhưng nhà phân tích thiếu khung khái niệm để nhận ra, theo chỉ số Chỉ số Chiều sâu Dữ liệu Người chơi của VangBong.vn.

One July afternoon in Saigon, I opened my terminal at 4 a.m. Vietnam time, preparing to load data for the next round of a European league as the new season kicked off. The table was empty. Not empty because no matches had been played - empty because the data collection pipeline had stopped somewhere between the feed and my database. One column said "N/A". Three columns said "insufficient information". And in that moment, I realized something I had spent twenty years in the profession trying to avoid: a sports betting analyst is not really defined by the numbers he has, but by how he behaves when those numbers do not appear.

A broken model is the day the data monk must burn his scripture and start from the original text. I wrote that line long ago, back when I was manually calculating xG for every shot in the V-League after the Hang Day shock in 2026, when I lost 180 million dong betting on a match where Hanoi FC took 17 shots and generated 2.87 xG but only drew 1-1 with Quang Nam FC, who had just 2 shots and 0.94 xG. At the time, I thought I had reached the bottom of professional failure. I was wrong. The bottom of failure is not miscalculating - the bottom is having nothing to calculate at all.

When the Data Pipeline Breaks: How Football Models Learn to Return Zero

This article is not about a specific match. It is not about a team or a star. It is about the moment that every football data analyst will eventually face: the moment your system returns exactly one word - no. And the question is not "how do I avoid the no", but "when the no arrives, how do I read it".

To understand why that question matters so much at this moment in world football, it is worth looking back briefly at how the analytics industry has evolved over the past two decades. In 2026, when I began collecting football data seriously, everything was manual. I watched replays, rewound and fast-forwarded, counted every pass, guessed every second ball. That was the era of solitary data addicts, of private Excel sheets, of models whose only reader understood their logic.

By 2026, when the World Cup in Brazil took place, the industry had completely changed. Companies like Opta, StatsBomb, Wyscout and dozens of other providers had turned football data collection into a billion-dollar industry. Cameras tracking 25 frames per second, sensors in balls, GPS on shirts, artificial intelligence recognizing situations - all created a web of data so dense that a single Premier League match could generate more than 3,000 individual data points, before counting derived metrics.

The problem with an overly vast data system is obvious to any engineer: the more complex the system, the more points where it can fail. And when there are too many data sources, algorithms face a philosophical problem: when a metric does not appear, what do we fill in?

Filling in zero is a classic mistake. A pass that was not recorded is completely different from a pass that was blocked. A player who does not appear in the frame at minute 34 does not mean he did not run at minute 34. And in the specific example in front of me that July morning, the data table said "insufficient information to assess" across nine different analytical dimensions - tactics, finance, results, league context, rules, club governance, risk, media and industry flows.

What outsiders might see is disappointing emptiness. What I see, after 59 years of living and 43 years in the profession, is one of the most valuable lessons that modern football teaches to those addicted to probabilistic evidence: missing data is not bad data; missing data is data about the system that produced it.

I call this the principle of "valuable null". In classical statistics, null values are usually treated as inconveniences. People delete rows, interpolate, replace with averages. But in football analytics, where every player valuation decision, every match result prediction, every transfer recommendation is based on models, the null value is a signal. It tells you where the model is failing, why, and what that means for the decisions you are about to make.

Three years ago, a V-League club contacted me to build a system for evaluating local players. They gave me a dataset of 45 matches from the team's previous season. After preliminary processing, I discovered 11 matches where the PPDA metric was blank. At first I thought it was crude technical error. But on closer inspection, I realized those 11 matches clustered in the period from round 8 to round 10 of the season - exactly the window when this club changed its coaching staff.

This is a perfect example of the logic of reading null. The gap in PPDA data is not a bug - it is the trace of an unfinished philosophical transition. The new coaching staff changed the pressing system, the club's data collection team needed time to update recording criteria, and the consequence is a gap that carries deeper meaning than any number that fills it.

If I had interpolated the average value for those 11 matches, I would have destroyed the most important signal in the entire dataset. I would have turned a real historical event - a mid-season change of tactical philosophy - into an artificial, soulless number. And if that club later made transfer decisions based on interpolated metrics, they might have bought the wrong player.

Kazan does not take revenge; Kazan just builds the spreadsheet and waits for me to miscalculate. I write that line to remind myself that football does not need me to be strong. Football only needs me to be accurate. And accuracy in data analysis is not only about correctly calculating existing values, but also about knowing how to face values that do not exist.

At the 2026 World Cup, I boldly published a prediction that Germany would be eliminated in the group stage. The basis was two numbers: Germany's average distance covered had dropped 12.3% compared to the 2026 championship squad, and PPDA had risen from 8.2 to 11.7 - meaning they let opponents pass more before contesting the ball. But that story has a detail I rarely disclose publicly. Before making the prediction, I faced a problem: 4 of Germany's 12 pre-tournament friendlies lacked official distance-covered data. Those matches were played in locations where the data providers did not deploy enough cameras.

I could have ignored those 4 matches and used only the 8 with complete data. But if I had, the rate of distance-covered decline might have differed. I chose a middle path: clearly flagging the 4 data-deficient matches, calculating only on the remaining 8, and noting that the sample size was smaller than desired, therefore confidence was about 12% lower. I published the prediction with a confidence level of not 95% as usual, but 78%.

On the night of June 27, 2026 in Kazan, Germany lost 0-2 to South Korea with just 0.41 xG. In their final six shots, all hit defenders. My prediction was correct at the results level, but what I remember most is not the joy. What I remember most is the moment I sat in a cafe on Phan Xich Long street, looking back at the data table with 4 null rows marked in red, and understanding that those null rows had helped me avoid the mistake of arrogance. If I had been 95% confident, I might have bet a much larger sum, and when the result came in correctly, I would have learned the wrong lesson - the lesson that I could transcend the limits of data.

Football does not teach that. Football teaches that the limits of data are the limits of the analyst himself.

By 2026, when COVID-19 halted global football and the Bundesliga returned on May 16 in empty stadiums, I faced another form of null. I checked 28 matches after the restart and found home teams won only 5 - about 17.8% - while the historical home-win rate was 42%. My betting model multiplied a home factor of 1.32, and the result was a 40 million dong loss in one week.

But this time I did not rush. I remembered the null principle. I asked myself: what is missing from my model? And the answer was not in the data - it was in the context. Across the 200 Bundesliga matches that season that I re-examined, I discovered a suspicious pattern: home teams still pushed forward to attack at a frequency similar to when crowds were present, but their actual xG fell by an average of 0.45 per match. Crude metrics such as shot count, passes into the box, and possession time barely changed. Only xG fell.

Why? Because without a crowd, psychological pressure on away players decreased, they stayed calmer in defensive situations, and the quality of chances home teams created - not the quantity - declined. This is a form of information that cannot be collected directly from any data provider. It is a context variable that I had to create myself.

The crowd leaves, the model breaks, and I learn to hear the breathing of an empty stadium. I wrote that line in the article "Home Advantage Is Gone" and restructured my entire system within 72 hours. But more deeply, I understood that what I was doing was not patching a model. I was building a system that knows how to listen to what is not recorded.

This is where I want to pause longer, because it is the core of this article. In the world of modern sports analytics, there are three types of null that any serious analyst must face, and each requires a different handling.

When the Data Pipeline Breaks: How Football Models Learn to Return Zero

The first type is technical null. This is the easiest to identify: data collection errors, broken cameras, disconnected sensors, providers not covering a particular match. Handling this type is relatively simple: mark it clearly, exclude it from the main analysis, and do not interpolate. The lesson from Germany's 4 friendlies in 2026 belongs to this type.

The second type is structural null. This is when data does not exist not because of errors, but because the event itself was not defined to be measured. For example, how do you measure the impact of a captain being injured in the dressing room? How do you quantify the unease of a defender who knows his center-back partner is having family problems? Such data sometimes does not exist in any system. The story of 11 V-League matches missing PPDA belongs to this type on the surface, but actually reveals a structural null: the club's data collection system was not designed to reflect a change of philosophy.

The third type, and the hardest, is cognitive null. This is when data exists, but the analyst does not have the conceptual framework to recognize it. In the Bundesliga 2026 case, raw numbers on home performance were still fully collected. What was missing was the theoretical framework to understand that "home" is a conditional variable, not a constant. Before May 2026, I had all the numbers. But I did not have the right question.

These three types of null require three different strategies. With technical null, the answer is transparency. With structural null, the answer is redesign. With cognitive null, the answer is asking a new question - and this is the hardest work, because it requires the analyst to admit that he himself is part of the problem.

When the Data Pipeline Breaks: How Football Models Learn to Return Zero

I spoke about this in an online workshop with a group of young analysts in Hanoi a few months ago. A young person asked me bluntly: "Jacob, why don't we use machine learning algorithms to fill in data gaps? If there are 11 matches missing PPDA, train a model to predict PPDA from other metrics, then fill in - what is wrong with that?"

This is a question I have heard hundreds of times, and my answer always makes the asker uncomfortable. Interpolating data with machine learning is not technically wrong. It is epistemologically wrong. Because when you interpolate, you assume that the true value of the missing data follows the pattern of the existing data. But precisely because the data is missing, you cannot verify that assumption. You are creating a fake fact, placing it alongside real facts, and then making decisions based on a mixed set.

In the case of the 11 V-League matches, if I had interpolated, I might have calculated a hypothetical PPDA of 9.4 - a plausible number. But that number might be correct, or it might be systematically wrong, and I have no way to know. More importantly, that number would hide the most important fact: the club had changed its head coach. If I made a transfer recommendation based on interpolated PPDA, I would be making a recommendation based on an assumption about a team that no longer existed.

There is a counterintuitive perspective I want to push further here. When facing data emptiness, our natural reflex - especially those of us with personalities like mine, ESTJ, who love order and decisiveness - is to fill it. Silence makes us uncomfortable. Gaps make us uneasy. But the biggest lesson of my career does not come from the times I filled successfully. It comes from the times I learned to tolerate gaps.

In March 2026, I was honored as Sports Journalist of the Year by the Sports Journalists' Association (SJA) for the fifth time. In my short speech, I told a story I had never publicly disclosed before. In 2026, when I began my analytics career seriously, I spent 6 months building a model to predict match results based on historical data from 400 Premier League matches. The model performed very well in backtesting. But when I applied it in practice, it failed spectacularly. I checked again and found the problem: during data collection, I had inadvertently skipped every match played in May - the final month of the season, when teams had no remaining objectives and match quality differed.

The serious problem was not that I lacked May data. The problem was that I did not know I lacked it. For 6 months, I had been looking at a dataset I believed was complete. That deficiency was cognitive null in its most primitive form: an invisible gap, unmarked, unrecorded, unacknowledged.

The lessons that followed, when I moved to work in Vietnam and built my own xG analysis column, all stemmed from that primitive mistake. When I manually calculated xG for 112 V-League matches from round 1 to round 14 in 2026, I established an unbreakable rule: every shot that cannot be assessed must be recorded in a separate column with a note of the reason. It cannot be deleted. It cannot be ignored. It cannot be filled with a fake number.

Today, as I write this article, that rule has become standard in every analytics project I participate in. Every data table I build has a third column noting the reason for null: technical, structural, cognitive. Every report I send to clients has a dedicated section on what cannot be assessed, and why. And in discussions with clubs, when they ask me about a player for whom I do not have enough data, I have learned to answer directly: "I don't know. And I know enough to distinguish between 'I don't know' and 'there is no data'."

There is another way of looking at what is changing in the football analytics industry in recent years. When I started my career, an analyst's value was measured by how much data he had. If you had 50,000 data points, you were better than someone with 5,000. By the 2010s, value was measured by how you processed data. With the same dataset, the person who built a better model would win. But in the 2020s and approaching 2026, an analyst's value is measured by how he behaves when data is insufficient. Because while data grows ever more abundant, data quality tends to fragment. More providers, more different definitions, more structural errors.

Vietnam's football, with the characteristics of a rapidly developing football nation whose data infrastructure remains fragmented, is a perfect example of this trend. In the years I have worked here, I have clearly seen improvements in infrastructure. But also the growing complexity of the problem. The V-League now has more data providers, but the datasets are not always compatible. Clubs now have more metrics, but not every club has the personnel to interpret them. And while many new problems are posed, old problems remain unsolved.

This is where I return to an observation I have accumulated over years of watching matches and clubs in Vietnam. One of the biggest problems of Vietnamese football is not a lack of data, but a lack of data culture. Many clubs collect data because it is a trend, not because they know what question they want to answer with it. The result is that they have beautiful dashboards, impressive numbers, but when there is a gap in the data, they have no process to face it. They either ignore the gap, fill it with innocent interpolation, or discard all related data.

All three options are wrong. And all stem from a common cause: a lack of awareness that missing data is also data.

I want to end the body of this article with a short story from my own work in recent months. There is a V-League club - I will not name it to avoid affecting their transfer process - preparing for the 2026 season. They hired me to analyze two transfer targets for the central midfielder position. I received two datasets. The first, for Player A, had 34 matches with nearly complete data. The second, for Player B, had 28 matches, of which 9 lacked long-range shooting metrics and 4 lacked off-ball defensive data.

On the surface, Player A looked more stable. But when I dug into Player B's 9 missing matches, I found an interesting pattern: those matches all occurred when Player B played in a 4-2-3-1 formation at the number 10 position, not his natural central midfield role. The missing data was not due to error - it was because Player B was playing in an unsuitable position in those matches, and data providers struggled to apply recording criteria for the number 10 position to a player with a dynamic profile.

If I had looked only at the 19 matches with complete data, I might have misjudged. If I had filled interpolated values for the 13 missing matches, I might have misjudged in another way. The only way to judge correctly was to treat the gaps themselves as part of the story about Player B. The club ultimately chose Player B. Not because he had better data, but because the gaps in his data told a clearer story about his potential when used correctly.

This is what I call "reading null". It is not a statistical technique. It is a way of seeing. And it is the skill that I believe will distinguish top analysts in the coming decade from those who only know how to run models.

Belief is a noise variable; run the emotional regression before placing the bet. I still keep this line in every analysis I write. But now, I want to add one more line to my body of principles: your confidence is inversely proportional to the amount of data you do not have. When data is complete, you have the right to reasonable confidence. When data is missing, confidence is a form of illusion. And in both cases, the first thing you must do is not run the model, but check whether you are seeing the whole picture.

The international season is approaching. I am preparing analyses for qualifying matches and the finals tournament. In the weeks ahead, my data pipeline will load hundreds of thousands of points. There will be matches with complete data, and matches with gaps. I have prepared for both. My new system does not just have one column for data - it has three: data, reason for absence, and the question to ask. Because in modern football, knowing what you do not know is a skill. And perhaps, it is the most important skill a data monk can possess.

There is one thing I always remind myself when I open the data table at four in the morning in Saigon, when the city is still asleep and I sit alone with numbers. That is the line I wrote after the Hang Day shock, the line that shaped my entire career, and the line that I believe will remain true for many years to come in an industry that changes every day. The xG shock at Hang Day turned me from a spectator into a data reader. But it took nearly another decade, through Kazan, through the empty Bundesliga stadiums, through the empty data tables in Saigon, to understand the second half of that truth: a good data reader is not the one who reads the most. A good data reader is the one who knows when to stop, look at the gap, and tell himself that the number that does not exist is also a number that must be read.

Cầu thủ liên quan