Building AI that reflects Latin American realities
Language models are reshaping cultural creation, sharing and preservation, yet a persistent problem plagues many systems: when trained predominantly on English-language datasets, they reproduce skewed depictions of non-English-speaking regions. Developers who prioritise inclusive design can harness large language models to broaden participation, safeguard cultural heritage, amplify underrepresented voices and unlock novel creative possibilities. This vision underpins Latam-GPT, Chile's inaugural open source LLM built specifically around Latin American cultural contexts.
The National Center of Artificial Intelligence (CENIA) in Chile unveiled Latam-GPT in February 2026 at the Artificial Intelligence Action Summit in Paris, though work on the model commenced around 2023 when researchers began investigating how to construct a region-tailored LLM using existing open source foundations. According to Gabriela Arriagada Bruneau, Associate Researcher at CENIA and assistant professor in AI and data ethics at the Pontifical Catholic University of Chile, the team sought to "learn and build a platform that was open source and that paid attention specifically to the cultural language-based difference, but also the cultural history and points of view within the Latin American region".
Tackling the data representation problem
Latam-GPT does not position itself as a rival to ChatGPT, Google Gemini or comparable mainstream applications. Rather, it aspires to furnish a more precise and performant foundational layer for constructing region-specific AI tools. The motivation crystallised when CENIA researchers analysed the composition of widely-used LLMs. "We wanted to see what level of representation existed, and roughly under two percent of all the training data and training models we considered contained actual Spanish or Portuguese data", Arriagada Bruneau notes.
The investigation revealed complications extending beyond mere language coverage. A critical concern centred on how AI systems portray historical and cultural distinctions across Latin America. "It was not just about processing language, but also more profound questions," she explains. "How should a model refer to historical differences? We know that how you train the model and how you put any type of filtering or benchmarks is going to have an impact on how the output is going to represent those things."
To tackle these multifaceted challenges, the initiative emphasises partnership across more than 20 Latin American nations, drawing on universities, government bodies, civil society groups and grassroots communities.
Community-driven benchmarking and evaluation
The collaborative ethos extends to how Latam-GPT undergoes assessment. By constructing proprietary benchmarks, researchers can measure how accurately language models grasp region-specific knowledge. "We've built benchmarks, and I think this has been a very beautiful way of learning," Arriagada Bruneau reflects. "It has come up as a very nice community embrace of how we think about language models."
Choclo, an open source benchmark centred on Latin American cultural knowledge, exemplifies this approach. The tool encompasses more than 100,000 question entries spanning geography, wildlife, flora, customs, culinary traditions and notable individuals—capturing not merely facts retrievable from online sources, but knowledge contributed directly by community members.
Complementing Choclo is Trueque, another open source benchmark featuring 500 carefully selected questions designed to evaluate factual accuracy and cultural sensitivity. The team additionally created Copuchat, an experimental conversational interface that feeds training data into Latam-GPT.
Transparency, governance and ethical foundations
Openness constitutes a cornerstone of the endeavour, encompassing both open source code and open model weights. Parallel attention flows toward data stewardship and ethical AI considerations—spanning transparency, confidentiality, algorithmic fairness and linguistic inclusion. "You have what I would call 'traditional ethical principles' – transparency, responsibility, fairness, accounting for bias", Arriagada Bruneau states. "But we also considered, for example, criteria for linguistic diversity".
CENIA researchers have produced a transparency document outlining the rationale for data choices and privacy safeguards, developed alongside the GobLab UAI team. Given Chile's nascent AI regulatory landscape, much of the project's responsible AI framework rests on the team's own dedication to openness and ethical practice.
Scaling up: next steps and indigenous language inclusion
The second phase concentrates on enhancing Latam-GPT's reasoning and analytical strengths alongside broadening its training scope. CENIA is engaging anthropologists, sociologists and representatives from indigenous communities to examine pathways for incorporating indigenous languages and knowledge systems.
Strengthening ties with Brazil represents another priority, establishing a regional technical partnership that simultaneously challenges the notion that AI development must follow a single predetermined trajectory. As Arriagada Bruneau suggests, the deeper lesson may be this: "there are different ways in which we can make technology our own to serve the needs and the contexts of different people".


