Research Grants
The overall goal of the Lab’s grantmaking process is to support research that leverages AI to augment human tasks and abilities and to ensure that these collaborations are ethical and beneficial to society at large. To stay up to date on future calls for proposals, please follow the Lab's LinkedIn and subscribe to our email newsletter.
2025 Call for Proposals Projects
The Lab's 2025 Call for Proposals focused on how to design effective solutions for safe and ethical human-AI collaboration in real-world settings. The objective is to foster developments that leverage AI to augment human tasks and ensure that these collaborations are ethical, inclusive, and beneficial to society at large. As AI systems become increasingly sophisticated, there is a pressing need to understand and optimize how they can work alongside humans in various domains and to anticipate potential medium- to long-term impacts of such applications in critical sectors.
Information about the award winners is included below. Teams will have until December 31, 2025, to complete their projects.
Digital Afterlives? Mourning, Memory, and Grief Tech
Principal investigators: Joseph Davis (University of Virginia) and Micah E. Lott (Boston College); Notre Dame collaborator: Paul Scherz, Theology; additional team member: William Hasselberger (Catholic University of Portugal)
New AI tools—“ghostbots”—can analyze data from a specific deceased person, like text messages, emails, and videos, to create an interactive digital companion that simulates them. This project will examine the reasons behind the rise of ghostbots, the ethical and religious concerns raised by the technology, and the role of grief and mourning in human lives.
- Organized the Digital Afterlives: AI, Memory, Mourning conference, Universidade Catolica Portuguesa, October 2025
Enhancing Human-AI Collaboration and Policy in Emergency Response: Ethical Deployment of AI-Enabled Drones
Principal investigator: Ricardo Morales (Brown University); Notre Dame collaborator: Jane Cleland-Huang, Computer Science and Engineering; additional team members: Kaitlin Harris (US Air Force/SAF/AQRE); Demetrius Hernandez (University of Notre Dame), and Tristian Hernandez
The ability to integrate AI-enabled drones into emergency response operations offers the potential to significantly enhance the situational awareness and decision-making of emergency response teams. However, balancing AI autonomy with human oversight introduces complex ethical and operational challenges. This project aims to inform future Federal Aviation Administration regulations by developing ethical guidelines for the deployment of AI drones in emergency response.
- Navigating the black box: Operational lenses for AI-enabled drone governance, MIT Science Policy Review, August 2025
Digital Moral Twins: From Bioethical Principles to AI Ethics and Back Again
Principal investigator: Jeffrey P. Bishop (Saint Louis University); Notre Dame collaborator: Paul Scherz, Theology; additional team members: Emily Dumler-Winckler (Saint Louis University), Lydia Dugdale, MD (Columbia University Medical Center), Jason T. Eberl (Saint Louis University), S. Matthew Liao (New York University), and -Devan Stahl (Baylor University)
In the case of an incapacitated medical patient, surrogate decision-makers often struggle to predict the patient’s desired healthcare choices. This project seeks to evaluate one proposed solution to this problem: a personalized patient preference predictor (i.e. “P4”) AI technology. Drawing on the most recent advances in the field of bioethics, this project will also assess whether core bioethical principles can be applied to the emerging field of artificial intelligence.
Building AI Text Classifiers with Peacebuilders: A Human-AI Collaboration to Improve Conflict Analysis and Resolution
Principal investigator: Allan Cheboi (Build Up); Notre Dame collaborator: Lisa Schirch, Peace Studies; additional team members: Julie Hawke and Will O’Brien (University of Notre Dame)
By inviting peacebuilders into the process of designing and developing AI text classifier technologies, developers can not only increase these practitioners’ awareness of AI and willingness to integrate it into peacebuilding work, but also improve the quality and relevance of AI-generated classifications for peacebuilding around the world.
(Digital) Companionship in the Digital Age: On Human-AI Relationships and the Ethical Landscape Surrounding Artificial Others
Principal investigators: Robert Clowes (NOVA University of Lisbon) and Kesavan Thanagopal (University of Notre Dame); Notre Dame collaborator: Diego Gómez-Zará, Computer Science
AI companion apps like Replika, Character.AI, and Kuki allow users to create “artificial others” that can provide conversation, emotional support, and judgment-free interactions. Are the simulated responses of AI companions ethically problematic? Can AI companions be viewed as “persons” in a philosophical sense? And what are the ethical responsibilities of app developers when releasing such technology into society?
- Organized the 1st Symposium on Generative Companionship in the Digital Age: On Human-AI Relationships and the Ethical Landscape Surrounding Artificial Others, University of Twente, July 1-3, 2025
Image Descriptions Are Less Reliable Than They Appear: Support for Blind Users Assessing Capabilities of AI-Powered Access Technology
Principal investigator: Amy Pavel (University of Texas at Austin); Notre Dame collaborator: Toby Li, Computer Science; additional team member: Meng Chen (University of Texas at Austin)
Millions of blind and sight-impaired people across the world now use AI technologies such as ‘vision language models’ to access visual information in their daily lives. But these models can produce errors that can go unnoticed by users, as they are difficult to identify without sight. How do we ensure the meaningful agency of those utilizing AI tools when direct verification of AI outputs is impractical or impossible?
- Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions, accepted to The 27th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS), October 2025
2024 Call for Proposals Projects
The focus of the 2024 CFP is “The Ethics of Large-Scale Models.” Large-scale models are drawing significant widespread public interest, driven by an increased focus on applications that leverage this technology, such as ChatGPT, LLaMa, DALL-E, or Midjourney. Large-scale models are artificial intelligence (AI) systems pre-trained on large datasets for general-purpose use, often adapted for a specific task using a fine-tuning process. Their adoption and impact are accelerating, driven by their potential in various contexts. With the increased adoption of large-scale models, multiple stakeholders have raised concerns regarding the ethical issues surrounding their design, development, and use.
Examples of areas for research and scholarship include, but are not limited to, the following:
- Society, Community, and Culture
- Governance and Policy
- Interdisciplinary (Developing trustworthy systems, design principles, human-AI collaboration, ROI)
Information about the award winners is included below.
Conflict Moderation: Implementing Bridge-Building Design in LLM Models
Awardees: Emillie de Keulenaar (University of Groningen); Notre Dame Faculty Collaborator: Lisa Schirch, Keough School of Global Affairs
This project looks at how large language models reproduce conflicts imbued in minority language training datasets, and then proposes frameworks for training and fine-tuning models for bridge-building applications. First, it examines how five closed and open-source LLMs — GPT 3.5 and 4, Llama2, BERT, Mistral and Gemini — respond to hundreds of historically, politically, socially or morally controversial questions, or "controversy prompts", in one majority language (English) and at least two minority languages with a conflict history (for example, Armenian and Azerbaijani, Arabic and Hebrew or other). By comparing results, I look at the discursive ways in which LLMs respond to controversy prompts and pinpoint where there is significant divergence, gaps and biases that obstruct the exchange of information, context and other elements necessary for dialogue across language groups. Then, I formulate a framework to retrain or fine-tune LLMs with additional datasets and bridge-building methods, which I formulate in consultation with peacebuilding and computation experts at the Notre Dame-IBM Technology Ethics Lab, the Council for Tech and Social Cohesion and the UN DPPA Innovation Cell. Retrained or fine-tuned models can be used for content moderation in search, recommendation and other applications, which, when queried with sensitive and conflict-prone keywords ("genocide in Gaza", "Nagorno Karabakh", etc.), may prioritise content that facilitates contextual, historical and other understanding across languages.
Contextualising AI Ethics in Higher Education: Comparing the Ethical Issues Raised by Large-Scale Models in Higher Education Across Countries and Subject Domains
Awardees: Wayne Holmes, Caroline Pelletier (Institut "Jožef Stefan"); Notre Dame Faculty Collaborator: Ying (Alison) Cheng, Psychology
Many Higher Education (HE) institutions have published policies on the use of large- scale models (LSMs), a set of Artificial Intelligence (AI) technologies, in teaching and learning. Such policies appear often to focus exclusively on academic integrity but rarely acknowledge variations in the meaning of ‘integrity’ across different cultural contexts and disciplines. In response, this study aims to examine more contextual understandings of LSM/AI ethics in HE. The project will begin with a systematic comparison of LSM policies across two national contexts in which they are used extensively (US/UK). This will be followed by interviews of teaching faculty to examine how those policies are interpreted and enacted across a range of HE subject domains, from arts to natural sciences. Finally, the outcomes of the policy review and interviews will inform a scalable, mixed-methods survey of faculty from the two countries and the varied subject domains, to reveal generalisable insights.
Cultural Context-Aware Question-Answering Systems: An Application to the Colombian Truth Commission Documents
Awardees: Luis Gabriel Moreno Sandoval (Pontificia Universidad Javeriana); Notre Dame Faculty Collaborators: Matthew Sisk and Anna Sokol, Lucy Family Institute for Data & Society, and Maria Prada Ramírez, Kroc Institute for International Peace Studies
The goal of this project is to create a Question-Answering System for the Colombian Truth Commission, the first-ever digital archive of the peace process. However, the archive contains a vast amount of data, making manual analysis techniques impractical. Therefore, the system is necessary to facilitate navigation and understanding of documents. This initiative ensures global access to the archive and explores digital approaches used in peace processes. It contributes valuable knowledge to improve future peacebuilding and conflict resolution efforts worldwide.
Engaging End Users in Surfacing Harmful Algorithmic Behaviors in Large-Scale AI Models
Awardees: Wesley Hanwen Deng, Motahhare Eslami, Ken Holstein, Jason Hong (Carnegie Mellon University); Notre Dame Faculty Collaborator: Toby Jia-Jun Li, Computer Science and Engineering
Traditional methods of testing AI models for harmful algorithmic behaviors, such as algorithm auditing, can fail to detect major issues given these methods’ reliance on small groups of AI experts. Recent research has shown that end-users, armed with their relevant cultural knowledge and lived experience, can surface harmful algorithmic behaviors that are overlooked by expert-led AI auditing and red teaming. However, there remains a notable absence of tools, guidelines, and processes to facilitate public participation in surfacing harmful algorithmic behaviors in large-scale AI models. To bridge this gap, we propose designing, developing, and evaluating a user-centered, interactive tool that effectively engages end-users in onboarding, exploring, reporting, and discussing the potential harmful AI behaviors exhibited in large-scale AI models. Our goal is to enhance meaningful public engagement to cultivate a responsible and ethical landscape for large-scale AI models.
Ethical Deployment of Generative AI systems in the Public Sector: A Practitioner's Playbook
Awardees: Dhanyashri Kamalakkannan, Shyam Krishnakumar, Titiksha Vashist (The Pranava Institute); Notre Dame Faculty Collaborator: Georgina Curto Rex
Large scale AI systems, particularly multimodal Large Language Models(LLMs), hold immense potential in transforming how governments regulate, deliver essential services, and interface with citizens. LLM-powered technology solutions are expected to play an important role in making public services more accessible, enhance and personalise public service delivery of goods such as education and health, improve hiring and personnel management, and make policy processes more participatory across governments. However, deployment of generative AI solutions in the public sector, whether to improve public service delivery, augment state capacity, or create new public goods, varies in its larger purpose from end-user or enterprise applications and give rise to substantially different ethical challenges when compared to its application in the private sector. Public deployments of AI need to be citizen-centric, and keep the public good at the core through enabling democratic values like trust, accountability, transparency, protection of rights and ensuring adequate public oversight mechanisms.This project seeks to create an ethical framework for public deployment of Generative AI which will translate ethical principles and existing global and national-level guidelines into a practical and accessible practitioner’s playbook which key decision-makers in government and companies can use as an ethical fitness check to mitigate potential harms before deployment of AI across the public sector.
Ethical LLM-based Approach to Improve Early Childhood Development in Children with Cancer in LMICs
Awardees: Horacio Márquez-González (Hospital Infantil de México Federico Gómez); Notre Dame Faculty Collaborator: Nitesh Chawla, Computer Science and Engineering, and Angélica Garcia Martínez, Lucy Family Institute for Data & Society
This project aims to develop a prototype integrating Large Language Models (LLM) and Automated Speech Recognition (ASR) technologies to resolve communication barriers between caregivers and teachers, inside and outside the National Institute of Pediatrics, Hospital Infantil de México Federico Gómez (HIMFG) in México City. This prototype will allow health workers and teachers to address critical early childhood development (ECD) dimensions in children with cancer in accessing community health, nutrition, education, and parental care programs. Success relies on actively engaging teachers and caregivers in assessing needs (phase 1) and ensuring continuous improvement (phase 2), fostering inclusivity, and impacting prioritized groups positively. This holistic approach to health, nutrition, education, and parenting care services is anticipated to improve social and economic development in population subgroups. The ethical implications of this technology will be explored, and we’ll discuss expansion to other countries/regions.
Generative AI and the Social Value of Artifacts: The Case for Saving Photo Morgues
Awardees: Kafui Attoh (CUNY School of Labor and Urban Studies), Jamie Kelly (Vassar College); Notre Dame Faculty Collaborator: Don Brower, Center for Research Computing
The rise of generative AI will increase the value of physical repositories of knowledge that are not subject to digital manipulation. Going forward, we will increasingly need to be able to verify claims made in the digital world by using non-digital evidence. In some domains, verification will be impossible. However, that is not true of every domain. Given this, we call for the preservation of physical artifacts and the creation of new archives as a safeguard against the potential flood of AI-generated disinformation. Using the case of newspaper photo morgues, we argue that this is especially important in hard cases where market imperatives and copyright diverge for the public’s new interest in preservation. With the support of the Notre Dame-IBM Technology Ethics Lab, we propose developing a white paper aimed at exploring these themes.
How LLMs Modulate our Collective Memory and its Ethical Implications
Awardees: Jasna Čurković Nimac (Catholic University of Croatia); Notre Dame Faculty Collaborator: Nuno Moniz, Notre Dame-IBM Technology Ethics Lab and Lucy Family Institute for Data & Society
This project will assess the impact of large languages models’ (LLMs) such as GPT on the formation of collective memory. In traditional terms, collective memory is a dynamic product of data selectively provided mainly from the institutions of memory (archives, museums, media, schools). However, AI changes how we access and use data in a public space. The main research of this project pivots around the social and educational uses of AI or how it reshapes the process of declarative memory-making and influences an ethically sustainable and impartial social framework for the construction of collective memory. Practically, how may our use of LLMs be shaping our collective memory concerning critical historical events? Our hypothesis suggests that the probabilistic matrix of tools such as GPT are prone to officialise the most represented narratives in their training data, opening troubling avenues of action in cognitive warfare, with a significant potential to shape our collective memory. To test this hypothesis, we will examine several controversial world events in different languages used by GPT to understand their differences concerning historical factuality and representation across languages and, if so, what are the practical and ethical implications in education, digital policy, and peacebuilding.
How Well Can GenAI Predict Human Behavior? Auditing State-of-the-Art Large Language Models for Fairness, Accuracy, Transparency, and Explainability (FATE)
Awardees: Jon Chun, Katherine Elkins (Kenyon College); Notre Dame Faculty Collaborator: Yong Suk Lee, Keough School of Global Affairs
This research project targets a pivotal issue at the intersection of technology and ethics: surfacing how Large Language Models (LLMs) reason in high-stakes decision-making over humans. Our central challenge is enhancing the explainability and transparency of opaque black-box LLMs and our specific use-case is predicting recidivism—a real-world application that influences sentencing, bail, and early release decision. To the best of our knowledge, this is the first study to integrate and contrast three different sources of ethical decision: human, statistical machine learning (ML), and LLMs. Methodologically, we propose a novel framework that combines state-of-the-art (SOTA) qualitative analyses of LLMs with SOTA quantitative performance of traditional statistical ML models. Additionally, we compare these two approaches with documented predictions by human experts. This multi-model human-AI approach aims to surface both faulty predictions across all three as well as correlate patterns of both valid and faulty reasoning by LLMs. This configuration offers a more comprehensive evaluation of their performance, fairness, and reliability essential for building trust in LLMs. The anticipated outcomes of our project include a test pipeline to analyze and identify discrepancies and edge cases in both predictions and the reasoning behind them. This pipeline includes automated API scripts, an array of simple to complex prompt engineering strategies, and well as various statistical analyses and visualizations. The pipeline architecture will be designed to generalize to other use cases and accommodate future models and prompt strategies to provide maximal reuse for the AI safety community and future studies. This project not only seeks to advance the field of XAI but also to foster a deeper understanding of how AI can be aligned with ethical principles. By highlighting the intricacies of AI decision-making in a context fraught with moral implications, we underscore the urgent need for models that are not only technologically advanced but also ethically sound and transparent.
Impact of Generative Artificial Intelligence - ChatGPT - on Higher Education in the Global South: Ethics and Sustainability
Awardees: Helen Titilola Olojede, Felix Kayode Olakulehin (National Open University of Nigeria); Notre Dame Faculty Collaborator: Nitesh Chawla, Computer Science and Engineering
The ubiquitous nature of artificial intelligence permeates all aspects of social life, and education is not exempted. Academic and information integrity, lack of diversity and bias, and privacy and security of learners are some ethical issues in using Generative AI (GenAI), such as ChatGPT, in education. This raises concerns about how educators should use GenAI to design instructional materials that facilitate authentic assessment? And in what ways can AI be employed to promote critical thinking? This study addresses the lack of ethical and effective ways to use GenAI, especially ChatGPT, in teaching, learning and assessment. Through intervention design, survey, Focused Group Discussion (FGD) and a series of structured and unstructured interviews, the study targets higher education lecturers, especially open and distance learning institutions, the major components of the Nigerian higher education sub-sector. The proposed project seeks to respond to the dearth of systematic training on applying GenAI in education among university lecturers in Nigeria, with the ultimate view of suggesting sustainable strategies for promoting ethical, effective and enduring use of GenAI across the Nigerian higher education sub- sector.
LLMs and a Well-Rounded Approach to Human Flourishing
Awardees: Avigail Ferdman (Technion-Israel Institute of Technology); Notre Dame Faculty Collaborator: Don Howard, Philosophy
Human flourishing ought to be an important ethical concern in Large-Scale Models, but it has yet to receive systematic scholarly attention. This research offers to address this gap, by combining perfectionism—an ethical approach to human flourishing—with an analysis of Large Language Model (LLM) environments. According to developmental perfectionism, humans flourish when they develop and exercise their capacities (to know, create, be sociable, exercise willpower) in well-rounded ways. Capacities are shaped by affordances—action possibilities in the environment. The research will analyze properties of LLM environments (e.g. content generation; speedy data analysis; user interface), according to their affordances (or constraints) for the competent exercise of a well-rounded combination of human capacities. This will provide a new lens from which to evaluate the goodness of LLM design, deployment and use, as well as an opportunity to offer an LLM ethics that goes beyond its current focus on risks and harms.
Mitigating Ethical Risks in Large Language Models through Localized Unlearning
Awardees: Alberto Blanco-Justicia, Josep Domingo-Ferrer, Najeeb Jebreel, David Sánchez (Universitat Rovira i Virgili); Notre Dame Faculty Collaborator: Nuno Moniz, Notre Dame-IBM Technology Ethics Lab and Lucy Family Institute for Data & Society
During training, large language models (LLMs) can memorize sensitive information or capture biased/harmful patterns present in their training data, which can then be delivered to end users at inference time. These undesirable behaviors undermine societal values and raise ethical risks. The overarching goal of our proposal is to develop an effective and efficient localized unlearning method that mitigates ethical risks in LLMs without compromising their utility. To achieve this goal, we plan to: i) precisely locate the minimal internal components of LLMs responsible for undesirable behaviors; ii) implement efficient target interventions on these components to unlearn those behaviors; and iii) evaluate our method using standard LLMs and data sets. Our expected outcome is to make LLMs more ethically compliant.
Research-Based Theater: An Innovative Method for Communicating and Co-Shaping AI Ethics Research & Development
Awardees: Anastasia Aritzi, Christoph Lütge, Franziska Poszler (Peter Löscher Chair of Business Ethics & Institute for Ethics in Artificial Intelligence, Technical University of Munich); Notre Dame Faculty Collaborator: Carys Kresny, Film, Television, and Theatre
This project aims to develop and implement an innovative and participatory methodology for teaching, research, and science communication in the field of artificial moral agents (AMAs). Within a hands-on research internship for students at TUM, the project team and students will conduct qualitative interviews on ChatGPT’s role as an AMA, corresponding societal implications and recommended system requirements. These findings will be translated into a theatre script and (immersive) performance. This performance seeks to effectively educate civil society on up-to-date research in an engaging manner and facilitate joint discussions (e.g., on preferred system requirements). The insights from these discussions, in turn, are intended to inform the scientific community and thereby, facilitate a human-centered/value-based development of AMAs. This project should serve as a proof of concept for innovative teaching, science communication and co-design in AI ethics research, thereby acting as the starting point for similar projects in the future.
Seeing the World through LLM-Colored Glasses - Detecting Biases and Deficiencies in Language Model Presentation of Underrepresented Topics
Awardees: Muhammad Ali, Ricardo Baeza-Yates, Shiran Dudy, Resmi Ramachandranpillai, Thulasi Tholeti (Northeastern University Institute for Experiential AI); Notre Dame Faculty Collaborator: Toby Jia-Jun Li, Computer Science and Engineering
Sources of information on the internet such as Wikipedia have exhibited long-standing disparities in representation across demographic dimensions. Women and gender non-conforming individuals, racial and ethnic minorities, and people from the Global South have all faced difficulties in finding members of their communities or related topics in resources that are considered “definitive” tools for information. The creation of large language models (LLMs) like ChatGPT has introduced a new mediator of information that is increasingly being used in education and beyond. This project will study differences in how LLMs present information about underrepresented topics. Through queries about public figures and geographic locations, we will build metrics to measure disparities in discovery rate, consistency, and sentiment of model responses. These metrics will allow us to better understand the real-world implications of adoption of LLM tools and how future access to information might be skewed or limited by this new technology.
Technology Transfer and Culture in Africa: Large Scale Models in Focus
Awardees: Catherine Botha, Franklyn Echeweodor, Anthony Isong, Edmund Ugar (University of Johannesburg); Notre Dame Faculty Collaborator: Jaimie Bleck, Political Science
The proposed project comprises a focused, multi-disciplinary investigation of how technology transfer impacts on culture in Africa in the context of large scale models. The project will yield three deliverables: one national workshop, one international conference held in South Africa and one journal special issue devoted to the theme. The impact of technology transfer on culture is an underexplored theme in the literature, and the impact of large scale models is only recently attracting much attention, but not from the perspective of technology transfer. We contend that the theme would benefit from a multi-disciplinary interrogation, to direct policy and law-making within the African context, as well as benefit makers of technologies. A carefully considered theoretical grounding to policy and other decision- making in the area of technology transfer and its impact on culture is, in our view, a first step in understanding this rich topos.
The Ethics of Using Large-Scale Models: Investigating Literacy Interventions for Generative AI
Awardees: Ranjit Singh, Emnet Tafesse (Data & Society Research Institute); Notre Dame Faculty Collaborator: Karla Badillo-Urquiola, Computer Science and Engineering
This project will explore literacy as a precondition for the ethical use of large-scale generative AI (GAI) models. We will investigate how literacy interventions for students and parents become a site for empirical ethics by building their capacity to handle novel concerns around the increasing use of AI in ordinary settings. We focus on two kinds of literacy efforts: (1) events for community college students that take a gamified approach to testing GAI models; and (2) surveys that document families’ anxieties and aspirations around GAI. While the first effort utilizes events as interventions that collect data on students’ concerns, the second effort collects data on the concerns of families that will, in turn, inform the design of new interventions. Our analysis will reflect on the nature of the critical thinking skills that these interventions produce, and how these skills mutually shape the ethics of using large-scale models.
The Influence of Virtual Avatar Race and Gender on Trust and Performance: Understanding How the Appearance of LLM-Enabled Avatars Influences Work in Virtual Reality
Awardees: Lisa van der Werff, Theo Lynn (Dublin City University); Notre Dame Faculty Collaborator: Timothy Hubbard, Management & Organization
This project investigates the interactions between individuals and Large Language Models (LLMs)-embodied virtual avatars within virtual reality (VR), focusing on the influence of avatar race and gender. As new technologies like LLMs, VR, and conversational virtual avatars converge, they redefine the future of work, enabling unique collaborations between humans and AI. This project aims to understand how workplace diversity within these technological advancements impact worker dynamics. Through a series of laboratory experiments, we will explore whether and how the race and gender of LLM-embodied avatars affect user interactions, trust levels, and task performance. The study leverages a trust lens to examine biases and differences in engagement with avatars, challenging existing notions of diversity and inclusion. By addressing ethical considerations and providing evidence-based insights, the project seeks to inform future design and policy decisions regarding the deployment of virtual avatars in professional settings, informing ethical integration of AI in the workplace.
2023 Call for Proposals Projects
The focus of the 2023 CFP is “Auditing AI.” As AI becomes more sophisticated, auditing it will only involve increasingly complicated ethical, social, and regulatory challenges. Dimensions that require auditing must be identified, agreed upon, and measured. AI auditors must be trained. Policies must be developed to govern the operations, credentialing, and impact of audits.
Potential areas for research and scholarship included the following:
- Scope of AI audits
- Regulatory frameworks for AI audits
- Methodologies for AI audits
- Skills for future AI auditors
- Teaching methodologies for AI audits
- How AI audits may impact various sectors and industries
- Suggested best practices for AI audits
- Adoption and deployment of AI audits
Information about the award winners is included below.
AI Audits for Who? Asian Perspectives on Rebuilding Public Trust via Community Ethics and Conflict Resolution Mechanisms
Awardees: Mark Findlay (Singapore Management University), Sharanya Shanmugam (Singapore Management University), Zhang Wenxi (Singapore Management University), Willow Wong (Singapore Management University)
The governance of artificial intelligence (AI) to mitigate societal and individual harm through ethics-by-design calls for equal attention to responsible data use before public trust can be conferred to AI technologies. Since trust is fundamentally rooted in community relationships, AI regulators seeking public acceptance toward AI innovation must attend to community-centric pathways to integrate data subjects’ voices in AI ethical decision-making. While traditional actuarial methods in financial audits can indicate a diverse range of evidence used to determine legal compliance, the researchers suggest that community interests and data subjects’ voices should not be absent in AI audit models. This research proposal will explore Singaporean (and Asian) perspectives on AI regulation to inform the motivations for using AI audits to rebuild public trust. Research analysis on the proposed scope and methodologies of AI audits will be followed by recommendations on the relevant skillsets for future AI auditors.
Algorithm Auditing, the Book
Awardee: Christian Sandvig (University of Michigan)
Would-be algorithm auditors presently have little guidance to begin learning about their research method. This project proposes to use an experimental writing process—the “book sprint”—to produce a short book about algorithm auditing. The book would be written by successful auditors and their legal advisors in a concise, accessible style, and published by a respected press. As a crossover title it would aim at both potential auditors (including algorithm designers, investigative journalists, and academic researchers) and others who wish to understand this area (including policymakers, regulators, and the general public). It will be grounded in the longstanding social scientific audit study literature, also known as correspondence studies or paired testing, giving it a distinctive voice. These foundational ideas will be updated and applied to contemporary systems by leading researchers in computing, employing real-world examples from cutting-edge contemporary audits.
Audit4SG: Toward an Ontology of AI Auditing for a Web Tool to Generate Customizable AI Auditing Methodologies for AI4SG
Awardees: Cheshta Arora (Independent Researcher), Debarun Sarkar (Independent Researcher)
This project aims to develop an ontology of AI auditing, which will be used to build an auditing web tool. The target users of the tool are external AI auditors of AI4SG (AI for Social Good), who will be able to generate customizable AI auditing methodologies. Along with the custom AI auditing methodology generated by the user, the tool will provide the user with a report card noting the pros and cons of the chosen AI auditing methodology. The web tool will be a proof of concept whose underlying ontology will contribute to the field of relational ethics in AI auditing that can account for diverse interests, multi-directional processes, multi-scalar networks of actors, institutions, data, algorithms, infrastructure, values, and knowledges. For heuristic purposes and based on the team’s domain expertise, the project will limit itself to three domains of AI4SG: economic empowerment, education, and equality and inclusion.
Best Practices for Communicating and Writing AI Audits
Awardee: John Gallagher (University of Illinois Urbana-Champaign)
This project aims to provide AI auditors with the discipline-specific training required to reach multiple expertise levels in their reporting, while navigating potential pitfalls of ambiguity. It will achieve this goal via the analysis and synthesis of one-on-one interviews conducted with 107 machine-learning scientists and researchers (118 interviews, over 86 recorded hours). This dataset, gathered for an unfunded project, contains unanalyzed responses to direct questions about communication with AI scientists, domain experts (non-AI), and the public. This large amount of process-based information requires trained human coders to identify best practices and themes. Drawing upon communication frameworks from the field of writing studies, these findings will be synthesized into training modules and disseminated to the academic community.
Building a Model of Participation of Children, Families, and Communities in AI Audits for Educational Services in Brazil
Awardees: Bruno Bioni (Data Privacy Brasil Research Association), Marina Garrote (Data Privacy Brasil Research Association), Marina Meira (Data Privacy Brasil Research Association), Júlia Mendoça (Data Privacy Brasil Research Association)
This project will concentrate on methodologies for meaningful participation of children, families, and communities in AI audits, focusing on educational technologies in Brazil. Both awareness and tools to promote participation are lacking and needed, given i) the penetration of AI-based solutions in schools in Brazil and ii) the current process to regulate AI nationally, which has so far paid little attention to this issue. The same reasons for concern also reveal the good timing for the project, with a community of technology and children's rights activists and scholars that has become increasingly more engaged in recent years. The second goal is to move forward in building a model for stakeholder participation geared toward this specific public. The scope is justified by a favorable context on three fronts: hits and misses in law-mandated participation models, a robust framework for child protection, and a problematic but timely process of AI regulation. The activities will consist of workshops and qualitative research to produce a roadmap proposal.
A Capability Approach to Ethics-Based Auditing in Medical AI
Awardees: Mark Graves (AI & Faith), Emanuele Ratti (University of Bristol)
Recently, it has been proposed to address the ethical challenges posed by AI tools by taking inspiration from auditing processes. This approach has been called ethics-based auditing (EBA), and it is based on an underlying conception of ethics that significantly draws from the recent principled turn of AI ethics, which is notoriously fraught with difficulties. This project proposes an alternative framework for EBA that is not based on AI principlism. In particular, it aims at conceptualizing EBA on the basis of the capability approach. Rather than checking for compliance to vague principles, EBAs should investigate the impact of AI tools to capabilities. The team formulates a preliminary characterization of capability-based EBA in medical AI. Deliverables will consist of a manuscript delineating the framework, a prototype AI tool that can be used for both internal and external auditing, and a conference paper.
Course Development for EU AI Act Certified Auditors
Awardee: Ryan Carrier (ForHumanity)
Independent Audit of AI Systems (IAAIS) provides a comprehensive risk framework in the same fundamental principles as Independent Financial Audit. ForHumanity is working towards building such an infrastructure of trust by enabling independent, third-party-assured compliance with the law. ForHumanity, a non-profit, public charity, ensures auditor/pre-auditor independence and anti-collusion principles across the ecosystem while training experts who conduct audit/pre-audit services while upholding the Code of Ethics and Professional Conduct. This project proposal is to develop and deliver an online 30-hour ForHumanity Certified Auditor Course (FHCA) on conformity assessment under the EU AI Act, built on the foundations of ForHumanity’s audit criteria.
Domain-Specific Legal Explainable Artificial Intelligence for AI Auditing
Awardees: Łukasz Górski (University of Warsaw), Shashishekar Ramakrishna (Research Advocacy and Management School of Intellectual Property Rights)
The aim of this project is to study the feasibility of developing a hybrid symbolic/subsymbolic explainable artificial intelligence system for law. The main assumption of this project is that the inclusion of legal domain knowledge in a system, something a symbolic AI excels at, would help to generate explanations that are of high use for the potential users of legal AI systems. In order to achieve this aim, a proof-of-concept system is to be developed and evaluated.
Expanding AI Audits to Include Instruments: Accountability, Measurements, and Data in Motion Capture Technology
Awardees: Abigail Jacobs (University of Michigan), Emanuel Moss (Intel Labs), Mona Sloane (New York University)
“Expanding AI Audits” proposes to develop and extend AI audit frameworks to include hardware and other instruments used to collect data used in AI and other data-driven algorithmic applications. The project extends AI audit frameworks to assess not only the outcomes of an AI system, but also to examine the assumptions on which those systems are based. Doing so enables auditors to assess the validity of an AI system, its appropriateness for use in specific contexts, and the conditions under which such assumptions may fail to produce safe and effective outcomes. The project will examine the use of motion capture technology as a mechanism for data collection and AI development and produce an expanded audit framework for hardware and other mechanisms. This framework will be accompanied by a workshop convening technologists, audit professionals, and regulators to disseminate the framework and findings as well as a series of public events.
A Framework for Auditing Organizations’ Responsible AI Maturity
Awardees: Ravit Dotan (University of Pittsburgh), Ilia Murtazashvili (University of Pittsburgh)
The team is creating a framework to audit the responsible AI maturity of organizations. The framework will evaluate organizations on three aspects: their knowledge of AI ethics, how well they embed AI ethics practices into their workflows, and what oversight structures they use for AI ethics. Beyond an auditing framework, deliverables will include empirically based strategies and tools for increasing organizations’ responsible AI maturity. This framework is a part of a larger project. First, the proposed framework is based on previous work that was done with the support of last year’s Notre Dame-IBM Tech Ethics Lab grant. Second, the framework will be developed in the context of a new lab at the Center for Governance and Markets at the University of Pittsburgh. In addition to developing this framework, the lab will apply it to the local tech ecosystem. Last, the lab will partner with other teams to support similar projects in other regions.
From AI Audit to Accountability: Understanding the Policy Perspectives Required for Accountability
Awardees: Charles Ikem (PolicyLab Africa), Jerry Monwuba (PolicyLab Africa), Kazeem Oguntade (PolicyLab Africa), Gideon Osadolor (PolicyLab Africa), Cornelius Udeh (PolicyLab Africa)
The most horrific tales about algorithmic injustice are tied to pre-mature AI deployments leading to social and benefits injustice. The AI audit ecosystem remains fragmented with tools and frameworks scattered and, in some cases, closed and gatekept. Given the increasingly visible policy developments mandating audits and the proliferation of algorithmic products, algorithmic audits are increasingly critical tools for holding vendors and operators accountable. But without an understanding of the drivers of AI audit relating to a particular sector, it is hard for meaningful policy to drive and hold operators accountable. The researchers propose to map the AI audit ecosystem trends and their relationship with societal and technological innovation and identify policy mechanisms, frameworks, and practical recommendations for policy developments and future research to advance the AI audits for accountability.
A Model AI Audit Process in the Pharmaceutical Industry
Awardees: Nick Bott (Takeda Corporation), John Chan (Takeda Corporation), Dustin Holloway (Takeda Corporation), Alejandra Parra-Orlandoni (Takeda Corporation), Tim Smith (Takeda Corporation)
Auditing methodologies for artificial intelligence (AI) within the life sciences will become essential as AI-driven algorithms and their integral data find new uses across the healthcare value chain, become more impactful in medical decisions, and are subject to more sophisticated regulation. The team aims to leverage Takeda’s experience implementing complex regulatory compliance to design an end-to-end AI audit process that will be generally applicable across the pharmaceutical industry. They will begin by aligning on standard ontologies for AI risk. In parallel, the team will design tools and software applications to embed ethics into the engineering lifecycle. Finally, they will engage a variety of stakeholders and experts to design a complete process map for an end-to-end AI audit that can be practically implemented in a healthcare company.
*Note: Takeda Corporation has elected not to receive a monetary award.
Open Source Audit Tooling (OAT)
Awardee: Briana Vecchione (Cornell University)
Despite growing recognition of the importance of AI audits, current solutions often fall short of auditors’ goals for thorough and accountable evaluation. This is partially due to the fact that the field is early in its development, and many of the basic terms—such as “auditing”—refer to a variety of aims and methods. This work aims to identify existing tools and resources used by auditors when analyzing AI systems by developing a taxonomy, distributing a survey, and hosting rounds of interviews with audit tool developers and/or practitioners about the tools they have developed and used. In doing this, the hope is to illuminate the landscape of existing tools and encourage solutions that allow for rigorous and accountable scrutiny of AI.
Process Audits for AI Bias: A Streamlined Framework for Independent Auditing of Algorithms
Awardees: Shea Brown (BABL AI), Khoa Lam (BABL AI), Benjamin Lange (BABL AI)
Although artificial intelligence (AI) has reshaped humanity in various positive ways, its potential for harm—such as bias—has sparked intense debates about how such technology can be effectively managed and governed. In recent years, independent auditing has been advocated as an accountability mechanism in various legal and industry frameworks, but its effectiveness remains questionable in the absence of consensus on auditing standards. The researchers aim to develop a standard methodology for the bias auditing of algorithmic systems, namely a process audit. They apply this auditing framework to derive audit criteria for an AI bias process audit for the case study of New York City’s Local Law 144, which requires annual independent impartial bias audits for automated employment decision tools starting in January 2023. Lastly, the team deploys audit criteria by conducting bias audits of relevant organizations and evaluates the process audit framework in practice using qualitative methods.
Trauma-Informed AI: Developing and Testing a Practical AI Audit Framework for Use in Social Services
Awardees: Suzanna Fay (The University of Queensland), Philip Gillingham (The University of Queensland), Paul Henman (The University of Queensland), Lyndal Sleep (The University of Queensland)
AI is increasingly being used in the delivery of social services. Offering opportunities for more efficient, effective, and personalized service delivery, AI can also generate greater problems, reinforcing disadvantage, generating trauma, or re-traumatizing service users. Conducted by a multi-disciplinary research team with extensive expertise in the intersection of social services and digital technology, this project seeks to co-design an innovative AI trauma-informed audit framework to assess the extent to which an AI’s decisions may generate new trauma or re-traumatize. It will be road-tested using multiple case studies of AI use in child/family services, DFV services, and social security/welfare payments.
Unpacking Algorithmic Infrastructures: Mapping the Data Supply Chain in the FinTech and Healthcare Industries in India
Awardees: Shweta Mohandas (The Centre for Internet & Society), Amrita Sengupta (The Centre for Internet & Society), Yatharth (The Centre for Internet & Society)
Large-scale adoption of AI systems across different sectors over the last few years has also foregrounded concerns around algorithmic bias, transparency, data privacy, and safety, within an overarching framework of ethics. Mechanisms to develop and deploy ethical AI systems have been a matter of debate, especially in India and across the majority world, primarily given the invisibility of algorithmic infrastructures that underlie the digital economy. Through a study of the data supply chain infrastructure in the financial services and healthcare industries in India, this project will aim to critically analyze the ethical frameworks that are adopted (or lack thereof) to develop and deploy AI systems in these sectors. Based on learnings from the study, it will offer an overview of best practices and an assessment framework that may inform efforts in auditing AI, and help develop robust data management and regulation practices in adherence with global but contextual ethical standards.
Users’ Trust in Human Resource AI Tools and AI Audits: How Can AI Audits Potentially Help Users of Human Resource AI Tools Gain More Trust in the Tools?
Awardee: Tina Lassiter (The University of Texas at Austin)
AI auditing is evolving rapidly and will become increasingly significant. Legislation, government agencies, and private companies are introducing a wide variety of AI audits for various sectors and industries. This study will focus on audits of AI tools for human resource (HR) decisions, particularly hiring decisions. Such decisions are highly consequential. While there has been extensive research done with regard to AI audits in this field in general, the fairness of HR AI tools, and the trust users have in the tools, there has been less focus on how AI audits could specifically increase the trust in such tools. By studying how users (applicants, employees, recruiters, HR managers) feel about different AI auditing models and in particular which parts of audits affect their trust towards AI tools, this study aims to gain deeper insights into the value and problems of AI auditing for users of AI in general.
UX Work as an Auditing Opportunity: Exploring the Role of Conversational UX Designers
Awardee: Elizabeth Rodwell (University of Houston)
User experience (UX) as a field has expanded rapidly, with many companies rushing to hire user experience professionals for the first time, and the intersection of usability and artificial intelligence becoming a significant part of the field. Therefore, UX researchers and designers are in a unique position to contribute to a more accountable AI and to develop tools and processes to audit this technology. This project focuses specifically on UX professionals working within a challenging subfield of conversational AI: voice assistants. It proposes a fieldwork-based collaboration within which the researcher analyzes how social inequities are reproduced in AI and establishes which methodologies UX can contribute to increasing oversight and accountability for businesses. UX has already laid the groundwork to intervene when a company’s products are working against the best interest of its users. But it must go further, in always considering ethics part of the user experience.
2022 Call for Proposals Projects
This year’s CFP focused on identifying and funding practical and applied interdisciplinary projects focused on at least one of six core themes: the ethics of scale, automation, identification, prediction, persuasion, and adoption.
Information about the award winners is included below.
SCALE
The Complete Picture Project
Awardees: Devangana Khokhar (Outsight International), Louis Potter (Outsight International), Denise Soesilo (Outsight International)
Final Deliverable – Project Report
The Complete Picture Project (CPP) addresses hidden and pervasive AI and Machine learning algorithmic biases by constructing complete test datasets that better represent the true diversity of human societies and communities.
The global data landscape disproportionately represents already empowered individuals and demographics. The reasons for this are varied but it means that AI and machine learning (ML) models are also biased when trained on incomplete data sets, which in turn amplify any biases that may already exist. This can have severe impacts on the lives of millions of people who are already subjected to under-representation and discrimination.
According to Forbes, the global AI-driven machine learning market will reach $20.83B in 2024. Low- and middle-income countries have already seen a rapid expansion in applications using this technology. The humanitarian and development sectors increasingly make use of machine learning models to reach beneficiaries faster, understanding needs better and make key decisions about the form and execution of life-saving programs. How do developers and users ensure that Artificial Intelligence (AI) algorithms serve all the members of a community equitably and fairly?
The Solution that CPP proposes is to use certified balanced datasets that can be applied to develop, test and validate inclusive AI models and to identify potential bias. CPP builds independent, broadly diverse, representative (of individuals and of contexts) datasets that can be applied to a variety of algorithms to address biases in all phases of the development process: 1) early design, 2) testing and 3) adaptation and adoption
InterpretMe 2.0: A Web Tool for Community-Centered Interpretation of Social Media Posts
Awardees: Siva Mathiyazhagan (Columbia University), Desmond Patton (Columbia University)
Social media posts can be incredibly difficult to understand because they can be highly contextualized and hyper-local in nature and misunderstandings have real consequences. The social media monitoring and interpretation processes play a role among many other systemically racially biased factors in compromising the lives and livelihoods of community members. InterpretMe aims to humanize social media content and help stakeholders learn about restorative alternatives to punitive action. InterpretMe 2.0 will be an e-guide for the prevention of misinterpretation and promote a holistic approach that enables stakeholders to explore people beyond social media posts. This holistic community-centered approach will enable law enforcement, the judicial system, and other reporting professionals to reduce punitive criminal justice responses associated with social media surveillance and prevent the use of social media for mass incarceration.
Solving Ethical Challenges in the Design of Open-Source Environments: Scaling Urban Mapping Models in View of the Locus Charter
Awardees: Monika Kuffer (University of Twente), Lorraine Oliveira (Independent Researcher), Julio Pedrassoli (MapBiomas Project)
Deliverables: Deprived areas in Low- to Middle-Income Countries (LMICs) are a big urban challenge that requires consistent and updated information about their living conditions. However, due to the high costs of very-high-resolution (VHR) imagery and their computational constraints, little research addresses the characterization of such areas and even fewer city-wide analysis. Considering these challenges, Oliveira developed an innovative unsupervised machine learning model to spatialize and capture intra-urban deprivation in São Paulo using solely open data sources. The results pointed out four types of deprived areas in the city and acknowledged the differences among them. Through the award, the project goal is to expand Oliveira’s research to a regional scale—the Metropolitan Region of São Paulo—and improve the level of automation of the model. In view of the Locus Charter principles, the team aim at delivering a workflow that incorporates the needs of LMIC policymakers and that facilitates the model comprehension and its application. Considering the inequalities exacerbated by the COVID-19 pandemic, the project has high societal relevance since it does not require field surveying and rapidly generates useful information for decision-makers. During the following months, the team will access, collect and process open-source datasets, including the spatial features developed with stakeholders. Then, they will develop, optimize, and assess the algorithm, dealing with possible transparency, safety, and responsibility issues.
AUTOMATION
Artificial Justice
Awardees: Halsey Burgund (MIT Open Documentary Lab), Sarah Newman (Harvard University), Jessica Silbey (Boston University)
What would it look like if the legal decisions made by the US Supreme Court were “handed down” not by a group of nine experienced justices, but rather by an algorithm? What could go wrong? And how is this any less arbitrary or “just” than the current US Supreme Courts interpretations of the Constitution?
We know that humans are biased, and that machines built by humans and trained on human-collected data are biased, too. And yet, predictive systems are increasingly prevalent. We’re now seeing advances in natural language processing and generation that even a few years ago seemed like science fiction. Using a language model trained on Supreme Court decisions, Artificial Justice explores the possibilities of such technological “advances.”
In a time when so many legal battles are both political and sharply divided, and the most highly trained legal minds in our country can vehemently disagree about the same set of facts, some may hope for a technological solution. But how could we trust that this technology would exhibit benevolence or fairness? Can we gain new insights into the shortcomings of our own legal system by exploring such questions through a technological lens? By investigating the current capabilities of AI-enhanced “decision-making,” what can we learn that might help us influence development in a prosocial way?
Algorithms already control much of what we consume online and they manipulate us in ways that are both ethically dubious and often inscrutable. From social media companies to politicians to foreign actors, the instances of AI-enabled influence-peddling are nearly endless. Artificial Justice will take shape as a creative multimedia work that shines a light on these everyday manipulations by applying them to the most significant legal decisions made in the US, including those from the past, the present, and even the future.
Developing Model Legislation for the Operationalization of Information Fiduciaries for AI Governance
Awardees: Josh Lee (ETPL.Asia), Lenon Ong (ETPL.Asia), Elizaveta Shesterneva (ETPL.Asia)
The idea of technology companies as information fiduciaries has been well-received by academics, legislators and technology executives alike. Facebook’s Mark Zuckerberg, for instance, has publicly acknowledged the “idea of [Facebook] having a fiduciary relationship with the people who use our services” as “intuitive” and consistent with Facebook’s “own self-image."
While a promising legal doctrine that has seen traction in the US and other jurisdictions, what remains less clear is the scope of duties imposed on information fiduciaries in fostering responsible innovation, specifically in the field of AI. This is especially when the extensiveness of such duties developed in United States legal literature is constrained by First Amendment considerations which do not apply elsewhere. It is also less clear how the recognition of information fiduciaries as a legal doctrine should best come about – through private law or regulation, or an interaction of the two, and how it could interact with upcoming AI regulation such as the recently proposed EU AI Act.
Balancing the utility which use of AI brings to society, we aim to identify an appropriate level of scope and extensiveness for information fiduciary duties in the context of responsible AI, and propose model legislation that jurisdictions can refer to.
The Ethical Radicals
Awardee: Freyja Van den Boom (Bournemouth University)
Final Deliverable – Project Report
Many industries are disrupted by the ongoing digitization and the insurance industry is no exception. With the help of big data and AI, insurers have been able to improve their risk assessment, claim management and provide consumers with more personalized insurance that meets their needs. Despite the benefits, the adoption of automated decision-making (ADM) processes also pose serious risks. Research shows the potential for insurers to discriminate unlawfully based on people's willingness to pay and/or unintentionally. The increased risk for proxy discrimination is a good example where a seemingly objective factor such as postcode or the color of a person's car turns out to be a proxy for otherwise protected characteristics such as race or gender.
There is still much uncertainty about what is ethical and what is lawful when it comes to ADM which may stifle the uptake of otherwise beneficial innovations.
The ethical radicals project therefore aims to test the boundaries what is legal and ethical and help clarify the grey areas for insurers to decide about their use of ADM.
As a contribution to help insurers ensure their use of algorithms remains lawful and ethical, this project will prototype a tool for insurers for internal assessment of their algorithms and design a high-level framework with issues to consider when deciding upon the development/adoption and use of automated decision making to ensure insurance practices remain lawful, ethical and in the best interest of consumers. The project consists of stakeholder interviews, literature analysis and workshops. If you are interested to contribute do not hesitate to contact the PI of the project. We look forward to presenting the results this summer.
Ethics Experiment on Designing Character for AI
Awardees: Charles Ikem (PolicyLab Africa), Sudha Jamthe (Stanford University)
Final Deliverable – Working Paper
Today AI is designed without any character basis with random behaviors and personality and no transparency on how it would behave in certain situations. When AI becomes an integral part of our lives and does useful functions like senior care or engage with people in serious situations, their behavior must be thoughtfully designed and grounded onto clear values defined as their character.
We propose an ethics experiment to collect data to test the hypothesis that adding character with AI ethical tenets will ensure transparency and trust of AI with users.
This will involve testing the personality traits of AI of an existing AI to check AI’s gender, humanizing features such as tone and inclusiveness in its engagement on AI confusion matrix for false positives. We’ll collect empirical data to show that the character tenets that drive personality traits are more trustworthy and fairer to users.
Our project will aim to 1) validate/invalidate our framework to engage the UX designer throughout the AI lifecycle to ensure that AI has character development, and that the AI remains ethical. 2) Establish evidence and theory that the current process of character development for AI is limited in that it does not ensure that AI is designed with ethical principles and the personality of the AI is biased with random genderizing and humanizing of the AI with data gathered without respect for privacy of users. The output will be a reference AIX framework for designers to use freely to build ethical AI.
Exploring Local Post-Hoc Explanation Methods in Tax-Related AI Systems
Awardees: Marco Almada (European University Institute), Błażej Kuźniacki (University of Amsterdam), Kamil Tylinski (Mishcon de Reya LLP)
Final Deliverable – Conference Paper
Final Deliverable – Journal Article
Final Deliverable – Conference
The project aims to answer the question of how to design AI systems in tax law which are capable of helping taxpayers understand the decisions of tax organs and thus avoid litigations such as the Dutch SyRI case. We want to achieve it by preparing prototypically developed explanation solutions for AI systems (XAI) that is used by tax administration for detection of tax fraud, risk profiling and auditing (selecting tax inspections). The XAI solution will be designed for the taxpayers, who, as subjects of decisions rendered fully or partially by AI system, are primarily concerned with “why” questions. Thus system’s behaviour must be interpreted in order to provide certainty with regards to relevant factors contributing to a particular outcome. For these stakeholders local post-hoc explanation methods that embody counterfactuals appear most suitable. We will present these XAI solutions to randomly chosen group of taxpayers for evaluation purposes via questionnaires. The most comprehensible XAI solution will be evaluated by them as the best in terms of their use in a fair and transparent way, thereby contributing to the use of responsible AI systems by tax administration.
The project aims to prompt the data science and tax law communities to strive together for designing ideal AI system for tax related tasks, i.e., the one that combines high explanatory capability with low knowledge-engineering effort and yet being highly accurate.
From Ethical Models to Good Systems: A Data Labeling Service for AI Ethics
Awardees: Andrew Brozek (Craftinity), Thomas Gilbert (Cornell Tech), Megan Welle (Daios)
Final Deliverable – White Paper
Technical approaches to AI ethics presently focus on encoding abstract ethical values directly into the models being trained. Less attention has been paid to context-specific norms and risks. For example, a self-driving car fleet that recognizes pedestrians and cars but not potholes would do enormous damage to roads even if the model is perfectly safe in an abstract sense. At present, we are unable to track the relationship between training data and the resulting behavior of a deployed AI system. We are developing a system that automatically monitors how training data will be relied on by the system to conduct particular activities. This will permit 1) real-time monitoring of model outputs; 2) recognition and correction of unethical AI system behavior; 3) feedback between 1) and 2) so that context-specific norms are sustained over time.
Our intended deliverable is a whitepaper that frames the prototype’s technical contributions as implementable, marketable, and significant for AI ethics. We intend this document to present a case study of automated vehicles and the types of harm present in it, with particular focus on visualizing route navigation for automated vehicle driving behaviors. It will address the feedback between the trained model and salient types of harm: How does the choice of features impact forms of model bias? How often should the model be retrained? How costly is data collection and (re)training? What level of performance is desirable for particular tasks?
Promoting Human Values in the Design, Development, and Policies of Brain-Machine Interfaces
Awardees: Margot Hanley (Cornell Tech), Helen Nissenbaum (Cornell Tech), Meg Young (Cornell Tech)
Brain machine interfaces (BMIs) are used to treat a range of cognitive and sensory motor conditions, including brain injuries and paralysis. With continued advances in neuroscience and machine learning, we are likely to see BMI capabilities and applications grow, both in their technical capabilities and in their reach. While BMIs provide important benefits, they also pose a set of pressing ethical challenges. At some point they will likely be able to access our intimate inner lives: our mental processes, decision making, and emotions—terrain within humans as yet inaccessible. With this access, however, comes fraught questions around autonomy, agency, and accountability.
This project seeks to support the policymakers who will need to respond to the coming wave of ethical concerns presented by BMIs. The research will consist of two strands of work: a literature review and semi-structured interviews. Our literature review will examine definitions of different BMIs for a policy audience, inform vignettes on how BMIs may mediate human experience, and synthesize work-to-date on threats that BMIs pose to human agency and autonomy across application domains. We will also conduct interviews with a broad range of stakeholders, including neuroscientists, technologists, ethicists, and policymakers. From these two strands of work, we will develop a white paper that identifies and defines the relevant technologies, highlights the domains where such technologies are likely to be applied, surfaces the key ethical
IDENTIFICATION
Comparative Analysis of Risks and Benefits of Digital Identification Systems in DRC, Gabon, Cameroon and Republic of Congo
Awardees: Divine Enkando (Data Rights Lab), Narcisse Mbunzama (Digital Security Group)
This study aims to conduct a comparative analysis of the risks and benefits of the different digital identification systems used in the DRC, Gabon, Cameroon and the Republic of Congo with an overview of the different technologies used, data rights, privacy, security and the laws in each of these countries.
Indeed, with the development of new digital identification technology and recent progress in artificial intelligence such as facial recognition tools, it has become almost possible to manage thousands of personal information in record time and to identify in a precise manner people just by having access to their data such as fingerprints, photos, videos, voices, etc.
Although these systems make positive contributions such as in the fight against terrorism, the identification of criminals, and in the management of the population, etc., in some authoritarian and non-democratic countries, these systems can be abused to stop dissenting voices, to identify opponents, to track activists and human rights defenders, etc.
Ethical Issues Associated With Pervasive Eye-Tracking
Awardees: Shaun Foster (Rochester Institute of Technology), Evan Selinger (Rochester Institute of Technology)
Our project, “Ethical Issues Associated With Pervasive Eye Tracking,” aims to raise public awareness of the dangers eye tracking poses in virtual reality, especially when Big Tech companies like Meta are heavily involved with designing and servicing the “metaverse.” We’ll create a video for the Notre Dame-IBM Ethics Lab to host for public dissemination that features the two grant principal investigators, Evan Selinger and Shaun Foster, having their eyes tracked in virtual reality as they explore different scenarios. The video will include narration that explains some of the assumptions companies might make if they possessed this eye tracking data. Since the footage will feature the principal investigators, the project avoids the privacy concerns that can arise when relying on volunteers. Furthermore, since the video will popularize assumptions made in the eye tracking literature, it will not advance eye tracking studies and doesn’t involve the ethical issues that pertain to human subjects. We’ll build the virtual reality application featured in the video using a customized virtual reality headset that’s configured to perform eye tracking functions.
We can’t commit in advance to specific VR scenarios. Potential ones include the following:
-
A picture viewing room: the user’s gaze is tracked to discuss assumptions about disclosing intimate information.
-
A shop: while users consider making purchases, their eyes are tracked to discuss assumptions about interest and intent.
-
Quiz questions: while users consider questions that vary in degrees of difficulty, their eye movements are tracked to discuss assumptions about task performance.
A Responsible Development Biometric Deployment Handbook
Awardees: James Eaton-Lee (Simprints), Alexandra Grigore (Simprints), Stephen Taylor (Simprints)
Biometrics are increasingly used in digital development contexts for enhancing effectiveness and efficiency. While there are some policies, tools, and frameworks specific to biometric technology, and many outstanding broad tools on Data Responsibility, many of these resources are either high level or generalist—or extremely academic. There is relatively little translational material aimed at biometrics specifically.
We would like to produce a handbook including a tool for assessing suitability and “right fit,” picking the right technology, understanding technical pre-requisites, assessing privacy impact, monitoring safety and effectiveness, providing a set of tools for generalist tech4dev or development practitioners to safely and effectively assess whether biometrics are a useful tool for their projects, ask the right questions to consider how to deploy them safely, roll out the right policies and procedures for governed, safe use, and incorporate them into their projects with the right level of oversight for ethical use.
As part of this piece of work we will engage a small cohort of INGOs and form a steering group for input and review—ensuring our work is well-matched with their needs.
PREDICTION
Explainable and Auditable AI in the Nexus of Climate Change and Food Security
Awardees: Catherine Kilelu (African Centre for Technology Studies), Winston Ojenge (African Centre for Technology Studies), Joel Onyango (African Centre for Technology Studies)
We propose a project that illuminates the core theme of ethics in machine learning-based PREDICTION. Since AI is fast-gaining ground in the continent, we intend to address the ethical limits of prediction; ethical frameworks for the use of predictive technologies; policy guidance for accountability and recourse with respect to predictions and predictive technologies.
We shall piggy-back the proposed study on our current studies at ACTS; studies that collect crop management data and yield, including weather variations, from small holder farmers and use machine learning to monitor how evidence-based climate change influence food yields within select staple-grain-growing areas of Kenya.
We shall combine desktop study with evidence-based experimentation to establish:
-
Knowledge of how such data is governed;
-
The most common data and algorithmic biases, and the errors which are due to such biases, in an African context
-
How existing tools perform in measuring the biases;
-
A map of which machine learning algorithms record least errors for which predictive scenarios;
-
A framework based on the above information, and a policy brief proposal.
Identifying Common Typologies of Harm in Forecasting Systems
Awardees: Nathaniel Raymond (Yale University), Bahman Rostami-Tabar (Cardiff University)
Forecasting plays a critical role in guiding decisions and developing business strategies in many organisations. Despite a considerable body of research and practice in the area of forecasting, the focus has largely been on the potential benefits of using forecasting and how indispensable its methods are; however, less (or arguably very little) has been contributed (both in research & practice) on potential harm caused by forecasting. The forecast design to forecast implementation life cycle is not yet generally agreed and described in the literature, but we hypothesize that, regardless of sector or decision, the life cycle is routinized predictably stable across forecast types and applications. The aim of this project is to investigate the typologies of harm and the mechanisms by which they may occur in the forecasting process, which can be generalized, identified, and modified when the life cycle is commonly described. The project produces a catalogue of where forecasting may cause harm and provides recommendations to address issues of potential harm in the forecasting process.
PERSUASION
An Audit for Children-Nudging: Games and Social Media
Awardees: Marianna Ganapini (Union College), Enrico Panai (ForHumanity)
Final Deliverable – Audit Framework
Deliverable: Audit framework for evaluating the ethical use of nudging AI technologies in gaming and social media aimed at children, including best practices and risk-mitigation strategies
Examining Dark Patterns in Apps Used by Adolescents
Awardee: Sundaraparipurnan Narayanan (Independent Researcher)
Final Deliverable – White Paper
Setting the Context: A nudge is a function of any attempt at influencing people's judgment, choice, or behaviour in a predictable way. Nudges tap into cognitive heuristics (“System 1” mechanism as referred to by Kahneman) to execute the influence.
While the ethics of nudging have been examined from the perspective of autonomy, rational choice and transparency, the impact on children (specifically adolescents) are not widely examined, more so in case of Dark Patterns.
Dark patterns: Deceptive User interface (UI) /User experience (UX) interactions including nudges that are non-transparent, constraining or limiting choices and/ or not in the best interest of the users. Essentially nudges that provide preferential weightage to one of the choices, choices are hidden from the user, inducing a false sense of urgency, hides information and limits user choices are negative nudges or dark patterns
Motivation: There are two key reasons for considering the adolescent age group for our study on nudges and/ or dark patterns. They are:
-
Adolescents are in their normative stages of decision-making developments and could be prone to influences by nudges.
-
Adolescents do not have an age-appropriate app category. App platforms have apps classified as children’s apps for age groups less than 12.
Project proposition: The research intends to focus on examining the dark patterns in popular apps (Android and IOS) in categories including Education, Gaming, Communication, and Social and Dating used by adolescents.
Human-Beneficial Decision-Making by Means of Augmented Reality Serious Gaming
Awardee: Ida Romana Helena Rust (University of Twente)
Final Deliverable – White Paper
I will explore how adding serious game (SG) elements to augmented reality (AR) can be used to stimulate self-awareness in our relation with smart (AI-infused) technologies with a human-machine interface.
The disappearing boundaries between the human and smart technologies may elicit that people will not be able to distinguish between choices that come from ‘the heart’ and choices that are sublimely imposed by these smart technologies.
I hypothesize that in order to make human-beneficial decisions while being part of a smart technological environment, one must first become aware of one’s relation with the environment; then develop an experiential understanding of one’s self, to consequently consciously decide how to act human-beneficially.
SG is an effective means to reach the cognition by means of the affects. The decision to combine SG with AR is based on the increasing popularity in various (professional) settings of the latter.
The methodology of this research is initially philosophical in nature, where I will use (post) phenomenological theory, applied to examples of SG in AR, to understand how the engagement with these games alters the experience of the user. I will provide a theoretical understanding of self-awareness in a smart technological environment and the experiential alterations within the individual as the result of one’s engagement with AR and SG. In the second half, I aim at the formulation of rudimentary guidelines on how ARSG can be developed to increase self-awareness and enable human-beneficial decision making.
A Manual of Ethical UX Design Principles
Awardees: Shyam Krishnakumar (Pranava Institute), Titiksha Vashist (Pranava Institute)
Final Deliverable – Practitioner’s Manual
Modern UX principles are limited in the sense that they can influence as well as predict user’s behaviour while on a digital platform. However, they do not take into account the user’s overall interface with the physical world or even the user’s mental, physical and emotional wellbeing in the digital realm. For example, many social media platforms are designed with Aza Raskin’s “infinite” scroll’ as a core feature. This one feature has been widely adopted by social media, e commerce, and OTT platforms. While vastly “improving” user experience on the platform, it has been one of the major factors behind social media overuse and addiction, leading to loneliness and depression. Raskin himself apologised to the public in 2019 stating he “designed the service to create the most seamless experience possible for users, but did not foresee the consequences.” This, and many design choices (gamification for instance) warrant the need for new ethical principles which serve as fundamentals to keep in mind while creating new digital experiences. As we move forward in the 21st Century, we are blurring the lines between the physical and digital worlds. It is therefore pertinent to reinvent user experience to aid and improve life both online or offline.
This project aims to understand which design choices promote dark patterns, and may have long-term, multi-sided harms baked into them. At a conceptual level, the fundamental challenge is to find the ethical line between persuasion and dark patterns, given that its widespread application has fundamentally changed user behaviour in its favour. We attempt to engage in multidisciplinary research to create a manual of ethical UX design principles which keeps the human at the centre, and takes into account not just metrics like performance, but also behavioural, cognitive and emotional wellbeing. We seek to engage deeply with research in the fields of design, cognitive science, social theory and psychology; and finally bring together a community of designers to apply these principles in real-world use-cases.
ADOPTION
Assessing Africa’s Policy Readiness Towards Responsible Artificial Intelligence
Awardee: Erick Otieno (Reallink Ltd.)
Final Deliverable – Policy Brief
Statistics show an estimated African population of 1,340,598,147 as demonstrated by (Worldometers, 2021). Out of this total estimate, the age bracket of 0-14 is estimated to be about two-fifths while the age bracket of 15-24 is estimated to be one-fifth of the total Africa population (Economic Commission for Africa, 2016). The data show a population that is important in terms of future Artificial Intelligence strategies especially when looking at the ethical Artificial Intelligence dimension. Sustainability is contextualized to mean the ability to have long-lasting residual Artificial Intelligence interventions that positively impact generations and generations to come. This is because the demography that will have the longest experience with Artificial Intelligence is the younger generation. It therefore follows that one of the questions emerging is whether Africa is ready in terms of policy infrastructure to dive into the world of Artificial Intelligence for the benefit of its population. With this in mind, understanding the policy intervention ecosystem would be an important undertaking even as the Artificial Intelligence interventions continue to be a platform to offer solutions to the African continent. There is a need for evidence that informs policy development and deployment strategies towards the successful development of responsible Artificial Intelligence that is regulated and is adaptable by the intended recipients. Consequently, this research will be an important contributor to the already existing conversation on responsible Artificial Intelligence both within Africa and beyond. This research will adopt exploratory research in attempting to address the research question.
Diagnosis and Mitigation of Bias from Latin America Towards the Construction of Tools and a Framework for Latin American Ethics in AI
Awardees: Luciana Benotti (Universidad Nacional de Córdoba), Beatriz Busaniche (Universidad de Buenos Aires), María Lucía Gonzalez Dominguez (Universidad Nacional de Córdoba)
The general goal of our project is to disponibilize, adapt and develop tools and frameworks for detecting, preventing and mitigating unwanted biases in Natural Language Processing applications. We will be focusing our work in word embeddings: a widely used but very opaque building block for many NLP models and applications. In parallel, we will develop a good practice guide based on Human Rights principles for local developers of natural language-based systems in Spanish. The tools developed in this project will help developers and non-technical stakeholders to evaluate, detect and mitigate unwanted biases in models and data, contributing to build a Latin American IA ethics.
We are an interdisciplinary team within Vía Libre Foundation with a background and long experience in social science and computer science research. To this we add our strength in integrating academic work with public policy advocacy, an excellent relationship with civil society in Latin America and the possibility of giving increasing visibility to the work in our field. In addition, our team has a strong focus on gender diversity.
We want Latin Americans NLP practitioners to be able to use the existing tools and frameworks to build more fair, accountable and transparent systems. Nowadays there is no go-to off-the-shelf tool that assists practitioners in assessing the systems they create: we intend to change this. Success for this project is the integration of ethics and fairness principles in the software industry development lifecycle.
Duty of Data Loyalty Model Legislation
Awardees: Woodrow Hartzog (Northeastern University), G.S. Hans (Vanderbilt University), Neil Richards (Washington University in St. Louis)
Current data privacy laws fail to stop companies from engaging in opportunistic, self-serving behavior at the expense of those who trust them with their data. A legal duty of loyalty would be a revolution in data privacy law, which is exactly what is needed to break the cycle of self-dealing that is ingrained into the current internet. Data collectors bound by this duty of loyalty would be obligated to act in the best interests of people who expose their data and online experiences, up to the extent of their exposure.
The team will draft model United States federal and state legislation that would impose a duty of data loyalty upon companies with respect to the human information they hold. The model legislation will prohibit information processors from designing digital tools and processing data in a way that conflicts with trusting parties’ best interests. It will also include setting rebuttable presumptions of disloyal activity, and a private right of action. The team will convene meetings in person and via videoconference with stakeholder groups to comment on the model legislation, to include academics, regulators, and members of the private sector. The team also will produce an explanatory white paper for legislators.
A Framework for Identification, Review, and Resolution of Ethical Issues in Healthcare Machine Learning Projects
Awardees: Jeremiah Fadugba (University of Ibadan), Pamela Kimeto (Kabarak University), Moses Thiga (Kabarak University)
The increase in the development of solutions for healthcare using Machine Learning (ML) continues to raise key ethical concerns in bioethics such as Beneficence, Non-maleficence, Autonomy and Justice. An additional concern is that of Explicability occasioned by the ‘black box’ nature of ML.
However, ML practitioners generally lack the capacity to identify and address these ethical issues in ML algorithm development, testing and deployment process. On the other hand, ethics review committee members drawn from the medical field also lack sufficient, if any, understanding of ML. They therefore lack sufficient capacity to identify and guide practitioners and researchers on ethical issues in ML.
This project therefore seeks to develop a framework for assessing and addressing ethical issues in Healthcare Machine Learning projects.
Increasing Venture Capital Investment in Ethical Tech
Awardees: Ravit Dotan (University of Pittsburgh), Leehe Skuler (Global Impact Tech Alliance – GITA)
The field of artificial intelligence (AI) is currently evolving faster than regulatory bodies can manage. Therefore, we look to venture capital as an influential stakeholder that can develop and enforce the ethical governance necessary to align the field with human-centric values.
To date, however, only a handful of VCs claim to consider tech ethics in their investment decisions and management. This is due, in part, to a lack of coordination over how to assess the ethical dimensions of AI ventures, as well as the absence of tools for external oversight.
Our team will address this barrier by producing a practical framework for VC stakeholders seeking to incorporate ethical AI criteria in investment strategies, accelerating the adoption of ethical AI standards across the tech industry.
Reversing the Mirror: Toward Ethical, Community-Centric Biometric Governance
Awardees: Hanson Hosein (HRH Media Group LLC), Shankar Narayan (Independent Researcher), Nandini Ranganathan (CETI, Portland State University)
Discussions of biometrics rarely include, let alone take as a starting point, the perspectives of communities that have historically been impacted by surveillance technologies. Those communities often lack the capacity and fluency to discuss the full range and rapid adoption of biometric technologies, yet are likely to be heavily impacted by them.
To address these challenges, the Reversing the Mirror project will create a multi-modal convening that will demonstrate a new model of engagement, moving beyond community “input” to relationship-based collaborations that recognize and account for structural barriers, and appropriately incentivize and resource diverse participants. The convening will consist of two intentionally structured and executed parts—a first part in which a diverse set of impacted community leaders will build a community-centric approach and related policy proposals for biometric governance; and a second in which other decision-makers, including lawmakers, regulators, and technologists, engage with these community-centric policies and consider implementation pathways.
The project will result in increased capacity among all stakeholders to engage with one another in the context of biometric governance as well as on broader issues of ethics in the technological space; new substantive ideas and implementation pathways for biometric governance; documentation of relevant scholarship; and a toolkit of takeaways for future convenings, among other outcomes. Ultimately, we hope to demonstrate that careful attention to power structures in the tech space—and appropriate interventions in response—can truly change this important conversation, making it more inclusive and reflective of BIPOC and other impacted communities.
A Roadmap for Ethical AI Standardization
Awardee: Christine Galvagna (Technical University of Munich)
Final Deliverable – Discussion Paper
Deliverables: White paper and website helping policymakers and civil society incorporate interdisciplinary expertise into standards-setting for AI
What Really Works? A Study of the Effectiveness of AI Ethical Risk-Mitigation Initiatives
Awardees: Ali Hasan (BABL AI), Ben Lange (BABL AI), Shea Brown (BABL AI)
Final Deliverable – Project Report
We propose to conduct an empirical study of the effectiveness of various AI ethical risk mitigation initiatives. In almost all industries, from banking and HR, to health care and edtech, leadership is waking up to the fact that the AI they are using could cause harm, and that this could be a significant risk for their organization. As such, large organizations are implementing a wide range of policies, initiatives, and governance changes in an attempt to be proactive and use AI responsibly. However, as there is little evidence yet as to which interventions truly work, and little regulatory guidance, these initiatives are often best guesses rather than established best practices.
We hope to provide some insight into what is working and what is not, as well as preliminary explanations as to why these interventions fail or succeed. Through desk research, interviews, and industry surveys, our research will lead to a framework for the effectiveness of ethical risk mitigation initiatives in organizations. The framework will identify successful interventions and connect them to the main features of institutions and the socio-technical settings, providing a list of initiatives that worked and why they did in a particular setting.