AGORA-KI: When AI moderates social media
To save resources, social networks increasingly rely on AI-driven systems for content moderation. These so-called social AI agents can independently evaluate content, make decisions about the visibility of posts, and interact with users. “Public debates are thus increasingly shaped by systems that are not themselves accountable,” says Dana Mahr, who leads the AGORA-KI project at ITAS.
The collaborative project, funded by the Federal Ministry of Research, Technology, and Space since June 2026, investigates this development. Researchers at ITAS work together with partners from the fields of social research, legal and policy advice, and technology development. The goal is to empirically assess the consequences of automated moderation and to develop verifiable rules for the responsible use of such systems.
Involvement of marginalized groups
To this end, the ITAS team conducts interviews with human moderators, developers, government officials, and people affected by hate speech or erroneous content removal. “We want to learn how automated moderation decisions are made, how they are perceived, and to whom those involved attribute responsibility for them,” explains Dana Mahr. People from marginalized communities – such as LGBTQIA+, BIPoC, and neurodivergent individuals – are deliberately included in this process. They are particularly often affected by incorrect moderation decisions and can thus contribute valuable experiences with hate speech, overblocking (the erroneous filtering of content), or a lack of appeal options.
Assessment tool for moderation systems
Building on this, AGORA-KI aims to bring together affected individuals, experts, and regulatory authorities to jointly develop requirements for future moderation systems that are compatible with democratic standards. In future workshops, the researchers will examine whether these requirements will remain valid as the technology continues to evolve.
All results will be incorporated into a structured assessment tool. This tool is the first to bring together what various stakeholders expect from automated moderation and to establish criteria for measuring whether a system meets these expectations. (July 16 2026)
Further links or documents:
- ITAS project website AGORA-KI
- Article The algorithm that hates for our own good by Dana Mahr

