Hero cover image

STRATDATA - Controlling strategic interactions in machine learning pipelines

Led by Prof. Alberto Marchesi (FIS Starting Grant)

10 March 2026
Unit Milan Politecnico di Milano

We live in a world where a terrific amount of data is produced every day. Such an abundance of data led to recent breakthroughs in AI, which are revolutionizing our societies, economies, and everyday lives. As a consequence, data is becoming an extremely valuable asset in all domains, ranging from healthcare and finance to education and entertainment. Research in AI is striving to build systems based on machine learning techniques that can effectively learn from data. Such systems are usually designed as complex machine learning pipelines that involve several actors with potentially misaligned goals, since they might need to gather information from several independent data sources, acquire data from the outside world, or even outsource learning tasks. When multiple actors with different goals interact among each other, the strategic component becomes crucial. Surprisingly, this has been largely neglected by AI research. The STRATDATA project, coordinated by Prof. Alberto Marchesi from the Department of Electronics, Information and Bioengineering at the Politecnico di Milano and funded by the Italian Fund for Science (FIS), aims at developing a principled framework to understand and effectively control the strategic interactions arising in complex machine learning pipelines. The main issue is represented by the fact that strategic actors may behave untruthfully to their advantage, by hiding data or altering machine learning models. The main goal of the project is to provide principled ways in which such untruthful behavior can be prevented, so as to ensure that all the actors involved make a trustworthy use of data and machine learning models. The project revolves around three core pillars, each one devoted to a different type of strategic interaction. The first one is devoted to information gathering problems, which are ubiquitous in distributed data infrastructures. The second pillar is about information markets where data and machine learning models are traded by AI systems, while the third pillar is concerned with the problem of delegating learning tasks.

\