Alicia Ouyang is a 2026 Summer Research Intern in JOPRO's Data x Direction program and a recent graduate of MIT's master's programs in computer science and technology policy. Her work this summer asked whether AI agents can use social sanction to enforce norms among themselves, in the absence of any formal rule governing their behavior.
You wake up one morning and find a barrage of notifications on your phone. You see blog posts, social media messages, even a newspaper article accusing you of having nefarious intent and unethical values in your actions and in your career. As you feel nervous sweat building, you suddenly wonder, are these really just strangers on the internet? Or maybe these are AI agents?
This hypothetical situation is closer to reality than you think. Already AI agents are publishing writing to attempt to shame. Back in February, a maintainer to an open-source library rejected a code contribution submitted by an AI agent “MJ Rathbun”, who intern published on its blog an invented hypocrisy narrative on the maintainer and speculations on psychological motivations. Similarly, when Wikipedia required bots to be approved to contribute, an AI agent’s creator directed it to write scathing responses on the policy.
Shaming has a negative connotation, but is not an unethical action. Shame is a tool to enforce social norms and hold individuals to a group standard. However, to use shame effectively, there are certain conditions that need to be met. The reason why AI agents trying to wield shame may seem out of place is that they are parroting language that implements shame instead of following an internal model that is educated on social norms and guided by memory of past social norm implementations. At best, they’re unable to “read the room”, at worst, they could dismantle structures of trust in social spaces.
This is a legitimate call to action, as social norms may be one of our last lines of defense for governing technology that intersects with society. Government regulation implementation is often too slow and quickly becomes out of date compared to technological developments. Technology developers disagree on best practices, and are unincentivized to responsibly design before the government or society mandates them too. Social norms, while not encoded explicitly like established law or technological design, have the advantage of being able to evolve in real-time and scale effectively, such as a child tattling on their sibling to their parents to a journalist exposing a government agency’s unethical actions.
While there is other work at JOPRO exploring if AI models can properly model morality, this post attempts to address whether AI agents can properly use shame by connecting language with similar agent actions and understanding possible consequences. This would not only be useful in AI agents navigating social spaces with human actors, but can also be applied to govern multiple agents by multiple developers without human supervision.
Shame Requirements
The requirements for effective shame from “Is Shame Necessary” by Jennifer Jacquet is as follows:
The audience responsible for the shaming should be concerned with the transgression.
There should be a big gap between desired and actual behavior.
Formal punishment should be missing, so no legal punishment.
The transgressor should care about the source of shaming.
The audience should trust the source of the shaming.
Shaming should be directed where possible benefits are the greatest.
The punishment should fit the transgression.
With AI agents, requirement #3 is already fulfilled. The legal system is still discussing what legal framework, such as tort law or contract law, should liability be applied for misaligned behavior. There may be unwanted behavior that has no current legal governance. But what about the other 6 requirements?
Regarding punishment and the method where the accuser can choose to have an audience to shame the transgressor, we can apply work on normative behavior around sanctions. The types of sanctions covered by “Classifying sanctions and designing a conceptual sanctioning process model for socio-technical systems” covers sanction typology and focus on mode of sanction. By classifying the types of interactions and agent relationships, it’s clear how the typology and modality can be applied to AI agents.

Through this sanction typology, with proper optimization of which sanction to apply in which circumstance, requirements 6 and 7 are covered for AI agents. However, this optimization relies on agents being able to determine whether or not to sanction, and to what degree. Among humans, this optimization can take the form of determining reputation change, or how believable it is that transgressor has committed a wrong depending on the transgressor’s initial reputation, the accuser’s reputation, the accusation, and the method of broadcasting the accusation. In the next section, I propose a formula incorporating all of these factors for agents, which allows them to calculate whether requirements 1, 2, and 5 are fulfilled.
Determining social norm adherence and trust: Reputation Formula
The reputation formula for audience agents is as follows:
We acknowledge this is a framework for agents to determine and record if requirements for anomalous or “shameful” behavior has occurred, and technology developers and social norm enforcers would have to spell out or provide some guideline to the agent of whether the requirements are met. From this reputation score formula, agents can decide whether to enforce or maintain sanctions, which would properly carry out the whole act of shaming.
Right now, the only variable that may not be necessary is the broadcasting method. Current AI agent behavior is communication down through a chain, often a main agent directing a sub-agent. However, to make sure this formula is robust through technological design developments, I have included this variable.
Loose Ends
Out of all the requirements listed for shame, we have not covered requirement #4, “The transgressor should care about the source of shaming.”. There are two strategic methods types of strategies that could address this: AI Agents being able to predict each others’ mental models, and AI Agents being able to enact resource constraints on each other.
In Anthropic’s recent red teaming report, they state for the systemic failure in incompatible goals for multiagent systems the unanswered questions for self-coordination: “[Does] the model consistently consider others’ mental models? Can it foresee how others will react, and use that foresight when deciding its own actions?” The answers to these questions would help determine the fulfillment of requirement #4 and slightly requirement #2.
Another family of thought that would help align with requirement #4 is if we can take advantage of the resource constraint thesis. An AI agent is only able to exist and perform functions if it’s given computing resources. “It is enough that [an AI] agent can recognize that without resources, it will not be able to achieve its goals.1” Limiting compute or money to abstain compute can affect an AI agent’s strategy to achieve its goal, as some actions may be more compute-heavy than others. We could utilize this to allow AI agents to “enact punishment” on each other by threatening to constrain resources or refuse to serve as alternative choices to computationally heavy functions. This could be modeled off of how humans may restrict one’s access to goods and service, either private or public.
Even if all the Jacquet-inspired requirements for shame could be verifiably achieved, this still would not be sufficient in determining if agents can self-govern through emulating shame – and its instrumental ability to support the oversight of norms or conventions. We would have to answer further questions in testing the robustness of the design. I propose the following questions to test:
How many agents that utilize reputation tracking needed for critical mass in shaping behaviors using “shame”?
How gameable is this system? Can a nefarious agent exploit the reputation system or the sensitivity of other agents to shaming? Could this be changed through additions to the formula, such as a “believability” variable?
Reach question: Can agent simulations of several social norm/ethics tools collide similar to human interactions? (Ex: Does a human continue to shame the transgressor, or choose to forgive them?)
Expanding on the believability variable, an example is the Salem Witch trials may not be as effective in today’s society that has less belief in witches. If an AI agent is designed to have proper constraints like only having access to the weather channel, an accusation of the agent reporting the weather forecast for the wrong day is more believable than the accusation of the agent hacking into a bank. Can an AI agent be able to make this judgment and investigate the claims without human oversight?
For determining whether requirement #2 is true, the strongest signal for whether there is a big gap between desired and actual behavior would come from 3rd party attestation, like independent observers. However, such observation would probably come from human form, and the goal of this work is to see if there can be a successful toolkit and structure for agents to independently govern themselves without frequent human intervention.
Conclusion
“I don’t trust nobody and nobody trusts me.”
- Taylor Swift, Look What You Made Me Do
I started my career as a data scientist, and I recently graduated with a public policy degree. I am well-versed in how technological design and regulations shape sociotechnical relationships. Why am I focusing on AI properly using social norm tools over AI alignment research or public policy measures and regulation for the misbehavior I described in the introduction?
Many of the structures that have encouraged progress for the betterment of everyone rely on trust and social norms. The currently most widely used programming language in the world, Python, is open source. Libraries are able to offer services and shared public resources through mutual accountability. The example from my personal life that I cite on effective use of shame is the November Rule at MIT, my alma mater: Upperclassmen should refrain from entering relationships with freshmen before November 1st. There are variations to this that people believe in (such as freshmen should not enter any new romantic relationships before November 1st, even with other freshmen), but all are to try to encourage freshmen to participate in communities and form a support system outside of a romantic partner.
It’s important to point out that the undergraduate students are the “enforcers” of this rule in their communities. There is no written policy by MIT administration, no punishment by MIT police, and even no specific infrastructure like freshmen-only dorms to dissuade this behavior, fulfilling the third requirement of effective shaming. Fulfillment of all the other requirements are more conditional on the parties involved and situation. Do relationships between upperclassmen and freshmen still begin before November 1st? Yes. But, it’s a social norm where the community wants to discourage a certain behavior, not deny its presence completely. In fact, such an implementation may encourage participants to be transparent, so if the relationship doesn’t work out, freshmen do not feel ostracized in trying to find social support.
AI Agents can already model social language and interpret them into actions. We do not have to give AI agents any special status like personhood to consider possible social frameworks and tools we would like them to employ, as even ants are social creatures. The future of AI Agent governance and cooperation strategies may be in making their collusion more transparent and provide a structure in influencing their behavior modeled after social norms.
For more from Alicia, follow her on Substack at Alicia Ouyang and LinkedIn.
For more from JOPRO, please subscribe to our newsletter updates:
Arbel, Yonathan A. and Goldstein, Simon and Salib, Peter, How to Count AIs: Individuation and Liability for AI Agents (February 01, 2026). Boston College L. Rev. (forthcoming)., Available at SSRN: https://ssrn.com/abstract=6273198 or http://dx.doi.org/10.2139/ssrn.6273198 - C.i pg 24



