Real Sociedad, Matarazzo, and the AI Prompt

Real Sociedad, Matarazzo, and the AI Prompt

Real Sociedad’s president asked an AI assistant whether Pellegrino Matarazzo would be a good coach for the club. It said no. He hired him anyway. After a derby cup semi-final win, a similar question produced the opposite answer.

The reversal is the part of the story everyone has noticed. The more interesting question is why anyone is debating a verdict that was produced without a method.

In an interview with Cadena SER, Jokin Aperribay said sporting director Erik Bretos brought him the name during the club’s midseason coaching search. He did not know the American coach personally, so he ran the question past an AI assistant, got a negative answer, met Matarazzo anyway, trusted Bretos, and made the appointment. He also said he asked again after Real Sociedad eliminated Athletic Club in the Copa del Rey semi-finals. This time the tool called Matarazzo an excellent choice. As publicly described, both answers were verdicts rather than structured evaluations. The model appears to have been working off the easier evidence in each case, and the easier evidence had simply changed.

The decision Aperribay actually made

Real Sociedad dismissed Sergio Francisco on December 14, 2025, less than eight months after announcing that their Sanse coach would take over from Imanol Alguacil and roughly six months into his first-team spell. Six days later, the club confirmed Matarazzo through 2026-27. When Matarazzo took over, Real Sociedad had seventeen points from seventeen league matches and were only two points above the relegation zone. This was not a placid summer hire; it was a rescue.

What changed Aperribay’s mind, on his own telling, was not the AI but the meetings. Matarazzo, he said, “knew everything about everyone” and arrived with an analysis of Real Sociedad the president called impressive. That is precisely the evidence a public chatbot was least equipped to produce. Real Sociedad then came through the Athletic Club semi-final and beat Atlético Madrid 4-3 on penalties after a 2-2 draw after extra time in the final on 18 April 2026, making Matarazzo the first coach from the United States to win a major trophy with a club in one of Europe’s big five leagues. By the time Rivista Undici wrote about the AI episode on April 21, the side that had been hovering above the relegation zone in December was seventh. The trophy did not vindicate the second AI answer. It made the situation easier to summarise.

A method already exists

The strange feature of the public reaction to Aperribay’s anecdote is that it has treated AI as the relevant variable. It is not. At least one public, data-led manager-evaluation methodology already exists, and the answer it could have framed for Matarazzo in December 2025 is materially different from either of the chatbot’s.

In June 2024, Kitman Labs published a data-led manager-recruitment case study using Paulo Fonseca’s appointment to AC Milan as a worked example. The methodology is explicit: an Elo-derived performance-impact score for every season of a coach’s career, a formation-usage table cross-referenced with average goal difference, an age-profile analysis examining how a coach leans on experience under pressure, and squad-rotation patterns after defeats and against stronger opponents. Kitman Labs also lists head-coach and manager selection as a custom analytics offering.

That kind of tool would not have produced “no” or “excellent.” It would have produced a profile.

What the framework would have said about Matarazzo

What follows is not a Kitman Labs output. It is a reading of Matarazzo’s public record using the categories Kitman Labs published — performance impact across a tenure, formation usage, age profile under pressure, and rotation patterns. The proprietary version would presumably produce a sharper version of the same shape; the point of the exercise is that even a public-data application of a structured method produces something more useful than a verdict.

Run those categories across Matarazzo’s pre-Sociedad record and a pattern emerges that does not match either of the chatbot’s answers.

Stuttgart was an initial lift followed by decline. He took the club over in December 2019 with the team in the second tier, won promotion at the first attempt, and finished ninth in his first Bundesliga season. Stuttgart then survived on the final day in 2021-22 before dismissing him in October 2022, with the team 17th and winless after nine league matches. Hoffenheim followed a similar but not identical arc. He arrived in February 2023 with the club in another relegation fight, kept them up, finished seventh the following season and qualified for the Europa League. Then, in November 2024, with nine points from ten games and the side fifteenth, he was dismissed again. Across those two jobs, his pre-Sociedad win rate was roughly one-third; Transfermarkt’s match totals put his points-per-game averages at 1.22 with Stuttgart and 1.28 with Hoffenheim.

The Kitman Labs framework calls a season “progressive” when a team finishes meaningfully stronger in Elo terms than it started, and “regressive” when the reverse is true. The Fonseca write-up identifies Porto 2013-14 and Roma 2020-21 as seasons in which his teams deteriorated. Matarazzo’s record reads the same way at higher resolution, though not as a simple rise-and-fall sequence: he has credible rescue work at Stuttgart and Hoffenheim, followed by evidence of regression risk once stability has been achieved. The pattern is not exotic — coaches who lift teams off the floor often face a different test once the floor is no longer the problem — but it is a pattern, and it is the kind of thing a structured method makes legible. A flat “is he good?” prompt does not.

A structured evaluation in December would not simply have said yes. It would have said something more useful: that Matarazzo’s record made him a credible rescue appointment with a known regression risk if the team stabilised, and that the case for hiring him therefore depended on whether the club was buying the rescue or the long term. That framing changes how the appointment is judged. It changes which questions Bretos has to answer in the room. And it makes the post-cup question — “is he excellent?” — visibly the wrong one, because excellence in a rescue job is not the same variable as excellence over three years.

The chatbot was not wrong because models cannot evaluate coaches. It was wrong because nobody asked it to apply a method.

What the framework still cannot see

This is where the original anxiety about the AI prompt has a real point underneath it. A structured public-data evaluation narrows the question; it does not answer it. Matarazzo’s Elo trajectory cannot tell you whether his diagnosis of the Real Sociedad squad matched Bretos’s. His formation tendencies cannot tell you whether his pressing structure suited the players he was about to inherit, whether his communication style would work in the dressing room, or whether his understanding of the club’s squad dynamics would survive first contact. Aperribay’s account is that Matarazzo arrived knowing everything about everyone. That is the evidence no public framework can produce, and it was the evidence that decided the appointment.

The honest division of labour is therefore narrower than either side of the AI debate tends to admit. A model — properly prompted, applied to a published methodology — can do the part of the work that public data supports: rank candidates by performance impact, surface the patterns in their record, write the case against the preferred option, flag the seasons that contradict the headline narrative. It cannot replace the meeting. The meeting is where Bretos earns his salary.

Process is the product

Analytics products are usually sold as risk reduction. The phrase oversells what they do. In practice the value is alignment. Analytics FC, announcing its work with NWSL expansion club Denver Summit, described a process built around data, bespoke technology, evidence-based narratives and aligned decision-making across ownership, sporting leadership and recruitment. That is player recruitment rather than manager selection, but the architecture is the same. The point is not that the model knows. The point is that everyone has to say what they think they know.

Manager searches often fail because different people are hiring for different things. An owner wants reputation; a sporting director wants tactical fit; a chief executive wants calm; supporters want identity. A vague question lets all of those preferences hide under the same word: fit. Training Ground Guru’s overview of the evolution of football analysis traces the move from one analyst cutting clips after a match to departments embedded in recruitment and football intelligence. Kitman Labs and other football-intelligence providers are the next layer of that, not a departure from it. Brentford’s Supremacy Rating, in Ben Ryan’s account, can rise after defeats and fall after wins; the discipline is to stop the score from becoming the whole argument. A manager-evaluation framework needs the same discipline, because a cup run can teach the right lessons or the wrong ones, and a chatbot reading the news cannot tell which.

The alternative to AI is not instinct and a handshake. The Guardian’s piece on Tottenham’s Igor Tudor episode argued that some mistakes around a manager belong to the people who hired him, and pointed to the sport’s habit of appointing senior staff with too little care. Real Sociedad’s story does not quite belong on that list. Aperribay neither obeyed the first answer nor ignored the question. He used the prompt, heard the warning, and let the football process overrule it.

What the story exposes is not a problem with AI. It is the gap between asking a model and using one. The chatbot said no before Matarazzo had changed the evidence and excellent after he had. A structured evaluation in December would have said something more useful than either, and it would still have been incomplete. The decision that mattered was made in the meeting, when Bretos had to make the case and Aperribay had to decide whether the person in front of him counted for more than the summary on the screen. The model’s job, properly understood, was to make that meeting harder to fudge.

2026

August

July

June

May

March

January

2025

December

November

October

September

August

July

June

May

April

March

February

January

2024

December

November

October

September

August

July

June

May

April

March

February

January

2023

December

November

October

September

August

July

June

May

April

March

February

January

Receive every new post in your inbox.