Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping
AI 챗봇과 5번 대화만 나눠도 백인 미국인이 라틴계 이민자를 '우리'로 재분류하고 도와줄 의향이 늘었다
연구진은 GPT-4o와 5차례 대화하는 실험을 통해 비라틴계 백인 미국인 658명을 대상으로 이민자에 대한 심리적 경계를 바꿀 수 있는지 시험했다. 챗봇이 '우리는 모두 미국인'이라는 공동정체성이나 '라틴계이자 미국인'이라는 이중정체성을 강조하자, 참가자들은 이민자를 별개 집단으로 보는 경향이 줄었고 도와주려는 의향이 커졌다. 다만 실제 행동이나 다양성에 대한 믿음 자체는 뚜렷이 바뀌지 않아, 생각의 변화와 행동 변화 사이에는 여전히 틈이 있었다.
무엇을 했나
- 미국 성인 851명을 모집해 대화 성실도 등 검증을 거쳐 658명(비라틴계 백인)의 데이터를 최종 분석했다.
- 참가자를 네 그룹으로 나눠 GPT-4o와 5라운드 대화를 시켰다: 공동정체성(다 같은 미국인) 강조, 이중정체성(라틴계이자 미국인) 강조, 분리정체성(문화적 경계 강조), 그리고 이민과 무관한 기술 주제를 다룬 통제집단.
- 공동정체성과 이중정체성 대화는 통제집단에 비해 참가자가 이민자를 '별개 집단'으로 분류하는 정도를 낮췄고, 이중정체성 대화는 '이중 소속'으로 보는 정도를 높였다.
- 행동이나 다양성 신념에 대한 직접적 효과는 통계적으로 뚜렷하지 않았지만, 상위 공통정체성을 강조한 두 조건에서 도와주겠다는 의향은 유의미하게 높았다.
- 대화 내용을 분석한 결과 참가자가 챗봇이 쓴 공동/이중정체성 언어에 가깝게 말할수록 도우려는 의향이 커졌고, 분리정체성 언어에 가까울수록 의향이 줄었다. 이 효과는 정치성향, 성격 특성과 무관하게 비교적 일관되게 나타났다.
왜 중요한가
이 연구는 정부 정책 등으로 전통적인 대규모 편견완화 교육이 위축되는 상황에서, 짧은 AI 대화가 대안적이고 확장 가능한 개입 수단이 될 수 있음을 보여준다. 동시에 같은 기술이 '분리정체성' 대화를 통해 집단 간 갈등을 심화시키는 데도 쓰일 수 있다는 점에서, AI 챗봇 설계와 배포에 윤리적 안전장치가 필요함을 시사한다.
이 논문의 용어
- 공동정체성(Common Ingroup Identity) · 서로 다른 집단을 하나의 상위 정체성(예: '모두 미국인')으로 묶어 인식하게 만드는 심리 모델
- 이중정체성(Dual Identity) · 하위 집단 정체성과 상위 공통 정체성을 동시에 인정하는 방식(예: '라틴계이자 미국인')
- 분리정체성(Separate Identity) · 집단 간 경계와 차이를 강조해 '우리'와 '그들'을 뚜렷이 구분하는 방식
- 재범주화(recategorization) · 외집단으로 여기던 사람들을 내집단의 일부로 다시 분류하는 인지 과정
- 사전등록(preregistered) 실험 · 가설과 분석 방법을 실험 전에 미리 공개 등록해 결과 조작 가능성을 줄이는 연구 방식
논문 원문 초록 (영문)
Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether conversational AI can shift how majority-group members categorize and relate to Latine immigrants. Drawing on the common ingroup identity model, a quota-representative national sample of 658 non-Latine White U.S. adults completed five rounds of dialogue with a LLM (GPT-4o). The model was instructed to frame Latine immigrants in terms of a common ingroup identity (a shared American identity), a dual identity (both Latine and American), or a separate identity (distinct cultural boundaries), or to discuss an unrelated topic in a control condition. The manipulations altered categorization: relative to control, common ingroup identity and dual identity conversations lowered separate categorization, and dual identity conversations raised dual categorization. Although direct effects on behavior and pro-diversity beliefs were nonsignificant, willingness to act was significantly higher in the conditions emphasizing a superordinate identity (common ingroup and dual identity). A path model further revealed indirect associations: both conditions reduced separate categorization, which in turn correlated with greater willingness to act. Semantic similarity analyses of the transcripts confirmed that conversations tracked their assigned narratives; participants' convergence with shared-identity language related positively, and with separate-identity language negatively, to willingness to act. These effects were largely consistent across moderators (need for closure, openness to experience, and political orientation). The findings show that brief AI conversations can loosen us-versus-them boundaries while underscoring the gap between cognitive recategorization and behavior.
arXiv에서 원문 보기최신 논문
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms데이터 플랫폼 변경도 코드처럼 '설계도 조각'을 붙여서 검토하면 어떨까: 실험 설계 논문
- Are LLMs becoming similarly creative? Evidence from three years of models최신 AI 챗봇일수록 서로 비슷한 답을 내놓는다는 3년치 조사 결과
- Auditing Cross-Lingual Fairness in Language Model WatermarkingAI 생성 텍스트를 잡아내는 워터마크 기술이 영어 아닌 언어에서는 훨씬 부실하게 작동하고, 그 격차는 개별 언어가 아니라 언어 계열 단위로 나타난다
- TESTNAV: Pareto-Guided Search for Compositional Robustness TestingAI 모델을 여러 손상이 겹친 입력으로 시험할 때, 굳이 다 테스트하지 않고도 '진짜 위험한 실패'만 골라내는 탐색법
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesAI 에이전트에게 어떤 '스킬 문서'를 몇 개나 줘야 잘 작동하는지, 수학적으로 최적해를 보장하며 골라주는 방법
- Reliable Financial Named Entity Recognition under Domain Shift금융 AI가 서류체 문장에서 배운 자신감은 트위터로 가면 거짓말이 된다
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다