AddressHate detects hate speech, particularly its coded and disguised forms that existing tools miss. Our suite of tools enable institutions, policymakers, and digital platforms to stop digital hate from turning into real-world harm.
A world where hateful online discourse is stopped before it becomes real-world harm.
AddressHate is an independent nonprofit building technology to serve the public good. We are developing AI to detect harmful language online that current moderation tools miss.
AddressHate’s AI model detects hate more efficiently and accurately than existing models because it understands hate at the narrative level. We are creating superior moderation tools by translating leading academic institutions’ independent, peer-reviewed research into expert-labeled data to train detection models.
Rigor: Every finding is grounded in peer-reviewed research and methodological transparency.
Independence: Editorial and institutional independence from government, advocacy groups, political actors, and platforms.
Urgency: Digital hate causes real-world harm that demands immediate action.
Collaboration: The problem is too complex for any one discipline to solve alone.
Precision: We go deeper than keywords and slurs to catch what others miss.
A library of expert-labeled training data. Real online content and comments are labeled by experts against a comprehensive set of categories, which were developed by leading hate scholars and linguists to identify specific stereotypes that express antisemitic and other hate ideologies. This data is used to train our AddressHate AI model, train existing models (Claude, OpenAI) to improve their accuracy in detecting hate and highlight their current blind spots, and provide researchers in the field insights missed by current tools.
The AddressHate AI model. Technology designed to detect which narrative is present in online discourse. It is built using the library of expert-labeled training data (see above) and built by our machine learning scientists. This AI model can be used as a resource by partner organizations to detect implicit hate on social media platforms and AI model outputs for the purposes of furthering research in the space.
A way to hold accountable social platforms and AI companies. An independent measure of how social media and AI companies monitor hate and enforce their own safety policies.
Narrative tracking dashboards. Using our AddressHate AI model and library of expert-labeled training data, we surface how dangerous stereotypes take root and spread, so newsrooms, policymakers, and educators can act on it.
Better, easier, and less expensive data annotation. Annotating data means having subject-matter experts read real online content and flag hateful narratives so that an AI model can learn to recognize it. It is one of the most time- and cost-intensive parts of building an effective AI model for hate detection. We developed a tool that simplifies the labeling process for nonprofits, companies, and researchers tackling online hate.
The Digital Hate Review. The first academic journal devoted to digital hate, built to bring historians, linguists, sociologists, computer scientists, and policymakers into one field to address this interdisciplinary problem.
Although there is enormous investment in platform safety, the most dangerous discourse goes undetected. While the best platforms invest in moderation, they focus on explicit indicators of violence, child abuse, or harassment, but don’t have the expert knowledge to catch implicit signals in any one category. To have effective moderation, specificity is needed for each category.
We provide the necessary granularity and the implicit level of speech for one specific moderation category: hate.
A lot of hate tracking still runs off a list of banned words and phrases. You match posts against the list, and whatever matches gets counted. This method is simple and fast, but it treats the symptoms of the problem rather than the root cause.
The content that does the most damage does not use slurs at all. It leans on coded language that current moderation systems cannot catch. Our classifier is built to read the belief underneath a post, so it catches the coded and implicit material a dictionary can't see, and it keeps working as the language shifts.
We go a level deeper than explicit keywords. We look at language through the micro and macro lenses of subcategories, narratives, and rhetorics.
Each of these are a level of language: a belief system, a story, and individual words. The taxonomies defined by our research partners, define subcategories, narratives, and rhetorics based on the following:
A subcategory is the belief system. An example is that Jews control the banks, that they're loyal to each other before their own country, that the Holocaust gets exaggerated. Ideas like these have been around for centuries and do not change. What changes is the occasion someone finds to reach for them. These recurring beliefs are what our taxonomy calls subcategories.
A narrative is a belief aimed at something specific that just happened, like an election, a war, or a violent attack. The belief is already there; the event just gives someone a reason to reach for it and offer it as the explanation.
Rhetoric is the wording itself, the form the belief takes once it becomes an actual post: a meme, a joke, an emoji, a question that isn't really a question.
After a major attack makes the news, someone posts “worth asking who really benefits from this.” Read in isolation, that's just a question.
But within certain contextual digital landscapes, this language is referencing a specific, well-worn conspiracy trope. Attached to a recent attack, the same rhetorical statement becomes a suggestion that the event was staged by some powerful hidden hand. If the user exists in the antisemitic version of this trope, the user would understand the hidden hand to be Jewish. In other ideologies, it could refer to other generalizations.
None of that is explicitly stated.
The sentence is vague and “just asking questions,” which is what lets the writer gesture at the idea without ever stating anything explicitly conspiratorial. Because there is no slur or explicit language, a keyword filter couldn’t catch it. And yet the sentence is enough to complete the argument for the reader who already knows the trope.
Most tools only read the wording, so a post like that slips straight past them. Ours is built to recognize the belief underneath it, which is how it catches the same idea whether it shows up as a slur, meme, or question.
Right now we're building tools to track antisemitism, anti-Black racism, and queerphobia.
Over time, we want to help researchers, policymakers, and organizations make sense of every form of hate. The methods we're building aren't tied to any one kind of hate, so they carry over as we grow.
We monitor discourse on many platforms such as X, Reddit, 4Chan, Bluesky, Tumblr, Facebook, Instagram, YouTube, and 8chan. We are going to expand to LinkedIn, and many more soon.
We work closely with leading academic researchers and institutions. They created the underlying taxonomy and annotation standards that keeps our work rooted in real scholarship on hate and extremism.
The Digital Hate Review is the analysis we publish about online hate. It draws on the same data and classifications behind our work. It is the first academic journal devoted to digital hate, built to bring historians, linguists, sociologists, computer scientists, and policymakers into one field to address this interdisciplinary problem.
We help advocacy and monitoring groups measure the whole conversation with a rigorous, taxonomy-grounded, AI-scaled approach, instead of putting out curated incident reports.
We help for-profit moderation partners because we are driven by research and open about our methods that seek to understand hate, not just delete flagged text.
We help general-purpose AI and social media companies because our models are built and trained by experts, so they catch the coded and implicit hate that broad systems are set up to miss.
General-purpose models like ChatGPT are useful, and some groups now point them at this problem. But there are three key challenges with this approach. The first is that these models are meant for broad tasks. Second, these models change their back-ends often without alerting users. And the third, is that running moderation tests across millions of posts gets expensive fast. Many studies have had to shrink their data to afford it.
We are building our own classifier and training it on data our experts label by hand. So we can run the full corpus without a per-post bill, measure exactly where it's strong and weak, and improve the parts that matter most for coded hate, which is where off-the-shelf models are weakest. Every batch of expert labels makes our detection better, and that capability stays ours.
We build our models in-house and train them on data our own social scientists and researchers label by hand.
Instead of checking text against a list of slurs, the models learn the accounts, phrases, topics, and rhetorical moves that hate tends to travel with. That's what lets them catch coded and implicit content at scale, the kind a simple word match would miss, and keep up as the volume grows.
For antisemitism, we code against the published Decoding Antisemitism Lexicon (Becker, Troschke, Bolton & Chapelan, Eds., 2024, Palgrave Macmillan/Springer Nature) — six categories and 40 concepts, used as the book defines them. It is open access under CC BY 4.0, so anyone can read the source our labels are built on.
For the other forms of hate we track, where no comparable published framework exists, our own experts developed the taxonomy on the same pattern: a broad category, then the specific beliefs within it.
Grounding the models in clear, documented categories instead of a loose keyword list keeps our labels consistent over time and from one analyst to the next, lets outside experts check our work, and makes it possible to compare one kind of hate with another.
They're at the center of everything we do. Our social scientists and researchers write the definitions, label the training data, settle the hard calls, and keep an eye on the model's output for drift.
The AI takes their judgment and applies it across millions of items. It doesn't stand in for them, and every accuracy claim we make traces back to their work.
Knowing a post is hateful barely tells you what to do about it. What actually helps is knowing which belief it's spreading and how far that belief has already gotten. With that, you can hand a platform more than a takedown list. You can show them where an idea started, where it's headed, and where there's a real chance to stop it.
And these narratives aren't random. They tend to begin in the same corners, the smaller forums and communities where a belief feels at home before it goes anywhere. From there, a handful of accounts do most of the work of carrying it into bigger, more mainstream spaces. They also get harder along the way: a joke turns into a conspiracy theory, and the conspiracy theory eventually becomes a reason to act.
Watching that happen close to real time gives you a lot of openings. A platform can catch a belief slipping out of the fringe and into ordinary comment sections and step in before it reaches millions of people. An advocacy group can answer it with counter-messaging while it's still just an idea and not yet a settled worldview. Someone in policy can notice the same pattern climbing across several platforms in the weeks before a big event, and get ready instead of scrambling afterward.
None of that is possible if all you can see is that a post was offensive. It works because we read the belief sitting underneath.
Text is where we're strongest today, and we're working on images, memes, and video too, since so much coded hate now lives in pictures rather than words.
Both. We dig through large historical archives to set baselines and see the long-run trends, and the pipeline keeps pulling in and classifying new content as it appears, so the dashboard stays close to the current conversation.
We test the model against a gold-standard set our expert annotators have labeled, and we report precision, recall, and how often it agrees with expert judgment, not one tidy number.
Because coded hate depends so much on context, the bar we care about most is agreement with our experts. And we're honest about where the model does well and where it still needs a person to weigh in.
Expert agreement is simply how often the model's call matches what our trained annotators agree on. Those annotators are social scientists and subject-matter researchers who know each kind of hate.
We treat their consensus as the truth because coded hate is a judgment call, and it's one that general-purpose labelers tend to get wrong.
We run the same expert-labeled test set through the general-purpose models and compare them with ours side by side.
Off-the-shelf systems are tuned to catch obvious hate and to avoid false alarms, so they routinely miss coded language, dog whistles, and anything that turns on context, which is the exact material ours is built for. It's not that those models are broken. It's that hate detection is a small side feature for them and the whole job for us.
Context and intent are written right into our taxonomy and annotation guidelines, and our experts handle the tricky cases by hand.
We're trying to tell hate apart from criticism, reporting, counter-speech, and satire, including blunt political speech, so that we're measuring hate and not policing opinions. When a case is genuinely a toss-up, we'd rather send it to a person than flag it too quickly.
We work only with content that's already public. We don't touch private accounts, direct messages, or end-to-end-encrypted platforms, partly on principle and partly because our job is to make sense of public conversation, not to watch individual people.
We retrain as we add more expert-labeled data and as new coded terms and slang show up, and they show up fast, so the models have to keep pace. If accuracy starts slipping on our gold-standard set, that triggers an update too.
Researchers, policymakers, platforms, journalists, and civil-society groups — anyone who needs solid, structured intelligence on online hate. Really it's for people who have to understand a problem before they can act on it, and who need numbers they can stand behind.
We're working toward giving vetted partners direct API access to our classifications. If that sounds like a fit for your organization, get in touch.
Yes. We help reporters cover online hate with data, historical context, and expert analysis, so a story can rest on measured trends instead of a few screenshots.
A spike is where the work starts, not where it ends. From there, partners can dig into the narratives and the specific content behind the surge, brief the people who need to know, shape moderation or policy decisions, and then check whether anything they did actually moved the numbers.
The value is in the why and the what-next, not the alert on its own.
Our pipeline looks at public content across whole platforms rather than individual reports, so we don't have a public submission tool right now. If your organization has a specific research need, reach out to us directly.