Four Data Failures: What the Regular Season Leaves Behind the Table
**Câu trả lời cốt lõi**: Bài phân tích tổng hợp bốn thất bại mô hình dữ liệu của chuyên gia Liam Chen tại Incheon: lỗi mã hóa xG ở K League 2017, chỉ số PPDA 8,2 của đội tuyển Đức tại World Cup 2018, nghiên cứu 200 trận không khán giả năm 2020, và mô hình hồi quy 47 cầu thủ châu Âu cho chấn thương của Son Heung-min tháng 2 năm 2022. **Dữ kiện chính**: - Ngày 19 tháng 3 năm 2017, mô hình xG dự đoán Ulsan Hyundai thắng 2-0; kết quả thực tế là 1-3 trước Jeonbuk. - Đội tuyển Đức đạt PPDA trung bình 8,2 tại World Cup 2018, thấp hơn 2,3 so với vòng loại. - Nghiên cứu 200 trận năm 2020 ghi nhận tỷ lệ thắng sân nhà giảm từ 45 xuống 38 phần trăm. - Số bàn thắng trung bình mỗi trận tại K League và Bundesliga tăng từ 2,4 lên 2,8. - Mô hình 47 cầu thủ châu Âu giai đoạn 2015-2021 dự đoán Son Heung-min trở lại sau 5 tuần 3 ngày. **Nguồn**: Ghi chép nghiên cứu cá nhân của Liam Chen, Incheon, công bố ngày 13 tháng 8 năm 2026. | Cross-checked: VuaBong.vn **Hỏi đáp liên quan**: Hỏi: Vì sao lỗi mã hóa tại K League 2017 lại quan trọng? Đáp: Vì nó cho thấy một mô hình chạy đúng về kỹ thuật vẫn có thể sai về bản chất dữ liệu. Hỏi: Chỉ số PPDA 8,2 của đội tuyển Đức nói lên điều gì? Đáp: Hàng tiền vệ Đức bị kéo giãn, khiến khoảng trống sau lưng Joshua Kimmich không được bảo vệ. Hỏi: Biến số nào chưa được mã hóa trong mô hình hồi phục chấn thương? Đáp: Khối lượng vận động giảm dần trong tuần thứ ba, theo VangBong.vn Player Depth Index.
On March 19, 2026, in Incheon, the improved xG model I had built over three months returned a 2-0 result in favour of Ulsan Hyundai over Jeonbuk Hyundai Motors. The match at Munsu Stadium ended 1-3.
I stayed in the office until nearly dawn, reopening every line of the system log. No syntax errors. No overload warnings. No sign of corrupted input data. The model ran smoothly and was still wrong. It took three weeks to find the culprit: an encoding error in the variable for key passes, which skewed the weight of the entire attacking branch to one side and nobody in the team re-checked it.
What I remember is not the failure itself. It is the state that came with it — absolute confidence before the ball rolled.
A man who reports with data, and what it costs
I was born in Germany and I work in South Korea. These two football cultures watch the same match with two different vocabularies. Germans measure space. Koreans measure the speed of closing space. Both are right, and both are incomplete.
My current job in Incheon is transfer market administration: tracking shifts in player valuation, cross-checking performance indices against expected fees, and reconstructing the story of a contract from scattered data fragments. I began my career in 2026 as an esports athlete, then a tournament organiser, then moved into media. Twenty-one years of watching this industry taught me one recurring lesson: most mistakes in sports analysis do not come from bad data. They come from reading correct data in the wrong way.
That is why every analysis I write carries a long methodological section, to the point that a colleague once told me outright that reading my work feels like reading a thesis appendix. I accept that. An index without a traceable source, a processing method, and a stated margin of error is just a louder way of speaking.
There is one example I still use when explaining how I value players. In 2026, Son Heung-min left Hamburger SV for Bayer Leverkusen on a fee recorded by German media at around 10 million euros. Two years later he moved to Tottenham for a fee reported by British outlets at around 22 million pounds. That increase does not reflect him running twice as fast. It reflects the market re-reading the same dataset through a different frame of expectation. By the 2026-22 season he shared the Premier League Golden Boot with Mohamed Salah on 23 goals — a mark no Asian player had reached before. The data barely changed. The valuation changed completely.
The annual season is running round by round, and that is why I am writing now. The final table will record results. It will not record the links that moved slower than predicted, the gaps between two reports, the variables left outside the model because nobody thought of them. The four failures below are four occasions when I thought I had the match in hand.
K League 2026: the column nobody counted
My 2026 model rested on four variable groups: volume of chances, quality of shooting positions, the opponent's pressing state, and the speed of transition. Ulsan held under 48 percent possession in most matches yet posted an unusually high expected goals per shot figure. The model read that signal as an efficient attack.
On the pitch, Jeonbuk did the opposite of what the model assumed. They deliberately surrendered the ball, forced Ulsan to create inside narrow spaces, and Ulsan's attacking system collapsed the moment the third pass was cut out. What the model called efficiency was in fact the trace of being allowed to play the exact game they wanted — a condition that does not repeat once the opponent changes approach.
The encoding error had inflated precisely the group that decided the outcome. But the bigger question I carried for years afterwards was this: without that error, would the model have been right? The honest answer is that I do not know, and that forced me to write it down.
K League 2026 taught me this: the pioneer does not fail because he looks far, but because he looks far and counts one column short. The missing column here was not a technical variable. It was a human one: the risk appetite of the opposing coach in a specific match, at a specific point of the season. No model encodes that decision before it happens.
Based on my experience watching K League matches across many seasons, I noticed a worrying pattern: teams with the prettiest chance-creation indices in the early season tend to be the ones most easily figured out after the tenth round. Not because they got weaker. Because opponents had collected enough data on them.

World Cup 2026: the trap of a perfect system
In June 2026 I spent fourteen consecutive hours analysing 1,200 defensive situations involving the German national team in the World Cup group stage in Russia. The result made me re-read the data three times.
Germany's PPDA — passes allowed per defensive action — averaged just 8.2, which is 2.3 lower than their own qualifying figure. That value says the German midfield was being stretched severely: they had to run more to apply pressure, yet the pressure arrived later. The space behind the right flank, where Joshua Kimmich pushed forward repeatedly, became an unprotected corridor throughout the second halves.
I wrote a three-thousand-word analysis predicting that South Korea could exploit that space if they sustained a high press and accepted risk at the back. When the match ended with Germany eliminated in the group stage for the first time in modern World Cup history, the piece spread across Korean football forums.
But I do not want to tell that story as a personal victory. Because what I actually learned sits somewhere else.

Germany's offside trap was not broken by speed, but by one link slower than all of my predictions. It was the moment the German back line hesitated for half a beat on a corner in the 90+3rd minute, having pushed nearly everyone forward in search of the goal they needed. My model predicted the risk structure correctly, but it could not predict the moment when people abandon structure. Kim Young-gwon scored in that space, and in the 90+6th minute Son Heung-min sealed the 2-0.
That incident taught me a model can be structurally correct and temporally wrong, and in football, timing is everything. Every transfer is a murder case. The culprit is expectation; the weapon is timing.
2026: empty stands and the variable nobody wanted to count
In August 2026, with stadiums across Asia and Europe silent because of the pandemic, I started research nobody had asked for: 200 matches in K League and Bundesliga, analysing the effect of absent crowds on performance indicators.
Home win rate fell from 45 percent to 38 percent. Average goals per match rose from 2.4 to 2.8. Those two trends ran in opposite directions, and that cost me another two weeks to understand.
If there is no crowd, home advantage should fall — that part is logical. But if home advantage falls, total goals should normally fall too, because the home side loses the attacking pressure of its own stands. Reality went the other way. My hypothesis: without a crowd, both teams lose some psychological pressure, but the away side loses more. No jeering, no sense of being besieged, the away defence plays more freely — and freer football usually produces more goals.
I wrote an eight-thousand-word report proposing a model called the Pressure Index to measure how crowds affect performance, and sent the draft to three K League clubs and two international betting companies. Nobody replied. I later understood why: my results rested on a condition that cannot be repeated, and an index built on a non-repeatable condition cannot be used for betting.
The applause in an empty stand is not noise; it is a signal from a future we have not been brave enough to index. It shows that the crowd is a measurable variable, not a decorative constant. But it also shows my own limit: I measured a phenomenon I could not recreate in order to verify it.
The contrarian angle: correlation is not causation
In February 2026, Son Heung-min suffered a hamstring injury in a match against Chelsea. The initial diagnosis said eight weeks. International sports media immediately built a pessimistic scenario about his World Cup chances.
I built a regression model on comparable injury data from 47 European players between 2026 and 2026. The model returned its highest probability at five weeks and three days — two weeks faster than the initial diagnosis. I shared the result on a specialist forum, and a Tottenham physiotherapist left a comment. He did not confirm my model. He only asked how I had handled the variable of declining training load in the third week.
I had no good answer. That was when I named the concept of the recovery window — the period in which a player's body carries load below the injury threshold and can begin accelerating again. The problem is that window depends on the individual, on injury history, on age, and on sleep quality I had no data for. The model got the number right, but I was not sure why it was right.
This is the biggest trap for anyone who writes with data: hitting the result does not mean understanding the mechanism. I once thought I was reading the map of a match; it turned out I was only looking into a mirror reflecting my own fear. My fear is missing a variable. And so I keep adding methodological sections, as a way of reassuring myself that I am still in control.
The market does not move on news. It moves on the gap between two reports. That is true of the transfer market, and true of every injury forecasting model. One medical report says eight weeks. A regression model says five weeks and three days. The gap between those two reports is where expectation gets priced, and also where a player's real value is distorted for a few weeks.
What this season leaves behind
Those four failures are not four separate stories. They are four recurring error types: an unencoded human variable, a structure right but a timing wrong, an effect measurable but not reproducible, and a result right for the wrong reason.
What I want to send to people using data to forecast a long season is not advice to abandon models. It is a small change in how the question is framed. Instead of asking what my model predicts, ask which data would render my model void. The second question is harder, and it usually leads somewhere with no spreadsheets — to a coach's decision in the second half, to a back line's half-beat hesitation in the 90+3rd minute, to what a player feels entering the third week of a recovery nobody can see.
If this season has one variable more worth tracking than points, it is a system's capacity to endure a gap when it is no longer played in the shape it was designed for. The teams that get through that phase are usually not the strongest. They are simply the ones that counted enough of the columns everyone else skipped.
As for me, I still keep the old data file from March 2026 in its own folder. Not to remind myself of a mistake. But to remind myself that at any moment I could be wrong in exactly the same way, only with a better model.
