Skip to main content
AI Data

Partnering With Universities and Communities for Language Data

Short answer. The Puno Quechua corpus partnered with a university and a local community organisation separately, because they contribute different things and conflating them loses both…

Mumu D. · July 2026 · 11 min read

Download PDF

Short answer. The Puno Quechua corpus partnered with a university and a local community organisation separately, because they contribute different things and conflating them loses both. The four-phase participatory model runs planning, preparation, collection and deployment, with governance settled in phase two — before any data is collected, not after. Preparation set a CC0-1.0 licence, prepared seed content across agriculture, healthcare and technology, and localised the Mozilla Common Voice interface. Collection used voluntary skill-based contributions with community-led validation.


Communities for Language Data?

The Puno Quechua speech corpus offers the clearest published template I have found for this kind of partnership, and the detail worth noticing first is that there were two partners, not one.

The planning phase involved establishing partnerships with the National University of Altiplano Puno and the local community organisation Illariy Ch'aska, alongside identifying the ISO 639-3 code and assessing community needs.

A university and a community body, engaged separately, because they contribute different things. That distinction runs through every successful example in this space and it is the one most commercial projects collapse, usually by treating a university department as a proxy for community access.


The four-phase model

The Puno Quechua team structured their work as a participatory design process in four phases, and it is worth walking through because each phase contains decisions that are easy to defer and expensive to defer.

Planning. Identify the language precisely, including its ISO code. Establish the partnerships. Assess community needs. That last item is the one commercial projects skip, and it is what distinguishes a partnership from a supplier relationship.

Preparation. Set up data governance, in their case under a CC0-1.0 licence. Prepare seed sentences and questions, which for them covered agriculture, healthcare and technology. And localise the collection platform, in this case Mozilla Common Voice, into Puno Quechua.

Note the ordering: governance is decided in preparation, before any data exists. Deciding licensing after collection means renegotiating with everyone who contributed.

Collection. Voluntary, skill-based contributions across reading, speaking, listening and writing, with community-led validation and privacy-preserving processing.

Two things there. Skill-based contribution means people participate in the mode that suits them, which widens participation beyond those comfortable being recorded. And validation is community-led rather than external, which is both a quality decision and a governance one.

Deployment. Open release on Mozilla Data Collective, with certificates of contribution and voucher incentives.

The incentive model is worth pausing on. Certificates and vouchers rather than only cash, which reflects that in a community partnership contribution is partly reciprocal and partly recognised rather than purely transactional. That is not an argument against paying people, which I have written about at length elsewhere in this series. It is an observation that recognition has independent value in this specific model.


Why universities and communities are not interchangeable

Both bring things the other cannot, and the Tonalli Corpus consortium in Mexico illustrates the division of labour better than most because it names roles explicitly.

Their partner set includes an intercultural education university with deep understanding of indigenous languages and cultures in the region, providing insight into engaging with indigenous communities; a technical institution offering expertise in technology and engineering, critical in the data collection and processing phases; INALI, dedicated to preservation and promotion of indigenous languages in Mexico, ensuring linguistic accuracy and cultural sensitivity; and INPI, whose mission covers the rights and cultures of indigenous peoples.

The consortium also includes a musical group dedicated to strengthening Mexico's ancient languages, particularly Nahuatl, using songs as a mechanism for promoting preservation.

That last one is the detail I would draw attention to in a scoping conversation. A musical group is not an obvious partner for a corpus project, and it reaches speakers, contexts and registers that a university department does not.

What universities typically bring: methodological rigour, ethics review infrastructure, linguistic expertise in the specific language family, students who can be trained into the work, archival capacity and institutional continuity beyond a project cycle.

What community organisations typically bring: access, trust, knowledge of which varieties matter and to whom, judgement about what is appropriate to record and release, and the ability to validate content against lived experience rather than reference works.

A project with only the first produces methodologically sound data that the community had no say in. A project with only the second produces culturally grounded data that may not meet technical specification. The published successes have both.


The governance frameworks that now apply

This has moved from good practice to something closer to expected practice, and anyone commissioning this work should know the vocabulary.

FAIR and CARE principles are named alongside PILARS in the Language Data Commons of Australia's standards for longterm sustainability. FAIR covers findability, accessibility, interoperability and reusability. CARE covers collective benefit, authority to control, responsibility and ethics, and it exists specifically because FAIR alone can facilitate extraction.

Indigenous data sovereignty is the framing underneath. The GovLab's 2026 review states it directly: data sovereignty as self-determination, with communities seeking authority over how data about them is collected, used and shared. And a definitional point that matters for language work specifically: data includes culture, language, land and relationships, not just statistics.

Language is the data in these projects, which places them squarely inside the sovereignty conversation rather than adjacent to it.

The same review notes that Indigenous data governance is "integral to mutually beneficial research partnerships", and describes a process worth copying: a research team developed academic ethics resources and documents over several months, which were then reviewed by Indigenous group leaders, with the resulting materials treated as living documents that can be updated as applicable to other projects.

The recommended focus areas for future work in that review are a usable checklist: community and project context, the changing digital landscape, individual and collective knowledge protections, planned project outputs, and confidentiality and anonymity nuances.

That distinction between individual and collective knowledge protection is the one least familiar to commercial data teams. Standard consent frameworks handle individual rights. A song, a story or a ceremonial term may belong to a community rather than to the person who recorded it, and individual consent does not settle it.


Making partnerships last

Two examples show what durability looks like, and both point at structure rather than goodwill.

MILPA, the Mexican Indigenous Languages Promotion and Advocacy collective, is a partnership between faculty, graduate students and undergraduates at UC Santa Barbara and members of the diasporic Mexican Indigenous community in Santa Barbara and Ventura counties, with most community team members affiliated with the nonprofit MICOP. That collaboration has run since 2015.

CoEDL and AIATSIS in Australia established a partnership from the outset for analysis, documentation and archiving.

Two structural features stand out. There was a named liaison role, a specific research associate responsible for the relationship rather than it being everyone's responsibility. And the partnership worked in both directions: it improved access to collections for the research programme, and supported community groups to deposit materials with the archive, providing safekeeping for language material while making resources more accessible to Indigenous communities.

That reciprocity is the thing. A partnership where material flows one way is an extraction arrangement with a friendly name.


What is being built now

Two developments worth knowing because they change the landscape for anyone entering it.

The New Commons Indigenous Language Data Commons Incubator, announced in 2026, is a six-month capacitybuilding programme supporting Indigenous-led teams to develop data commons: community-governed datasets enabling responsible, equitable use for public-interest purposes. It was co-designed with a 19member global Steering Committee of Indigenous language data experts and implemented in partnership with the GovLab and Microsoft, providing mentorship, technical guidance and capacity building, with concept notes due in August 2026.

The significance is the model rather than the programme. Community-governed datasets is a different structure from either open release or commercial licensing, and it is being institutionalised with major backing.

Language data commons infrastructure is being built nationally in some jurisdictions. The Australian programme provides data governance frameworks respecting cultural protocols, support for securing vulnerable or at-risk language materials, help making materials accessible in appropriate ways, and tools and training enabling community-led use of language data.

For a commercial data operation, both developments point the same way: the infrastructure and the norms are being set by community and academic institutions, and the sensible commercial position is to work within them rather than around them.


Where we sit

Declaring the interest: Lifewood collects language data across 50-plus languages and dialects, and partnership-based collection is how low-resource language work gets done.

Three observations I would offer to anyone structuring this.

Partnerships take longer to establish than projects take to run. The MILPA collaboration is eleven years old. CoEDL and AIATSIS partnered from the outset of a multi-year centre. A commercial timeline that allocates three weeks to "establish community partnership" has misunderstood the unit of time involved. The practical implication is to build relationships ahead of demand rather than in response to a signed contract.

Name a liaison. The CoEDL model of a specific person responsible for the institutional relationship is worth copying exactly. Relationships that are everybody's responsibility are nobody's, and they decay quietly between projects.

Decide the governance before the collection, and be honest about the commercial position. A community deciding whether to work with a commercial data company is entitled to know what happens to the data, who can license it, whether it is exclusive, and what the community retains. Those are answerable questions and the answers may be less generous than an academic open-release model. Saying so plainly is better than discovering the mismatch after collection, and in my experience communities are considerably more willing to work with a clear commercial proposition than with an ambiguous one.


A partnership checklist

Engage universities and community bodies separately, and map what each contributes before approaching either.

Identify the language precisely, including ISO code and target varieties, in the planning phase.

Assess community needs, and be prepared for the answer to change the scope.

Decide licensing and governance in preparation, before collection.

Localise the collection platform into the target language rather than running it in a lingua franca.

Offer skill-based participation across reading, speaking, listening and writing, so people contribute in the mode that suits them.

Make validation community-led where the content is cultural.

Address collective as well as individual knowledge rights, since consent from a speaker does not settle community ownership of what they said.

Build reciprocity into the structure, so material and capability flow both ways.

Name a liaison and fund the relationship between projects, not only during them.

Reference FAIR, CARE and Indigenous data sovereignty explicitly in the agreement, because those frameworks are now the expected vocabulary.


Key takeaways

  • The Puno Quechua corpus established partnerships with both a university and a local community organisation, engaged separately because they contribute different things.
  • The four-phase participatory model runs planning, preparation, collection and deployment, with governance decided in phase two before any data exists.
  • Preparation included setting a CC0-1.0 licence, preparing seed content across agriculture, healthcare and technology, and localising the Mozilla Common Voice platform into the language.
  • Collection used voluntary skill-based contributions across reading, speaking, listening and writing, with communityled validation.
  • Deployment used open release with certificates of contribution and voucher incentives, recognising contribution as well as compensating it.
  • The Tonalli Corpus consortium in Mexico names distinct partner roles: intercultural education expertise, technical and engineering capacity, national language institutes for linguistic accuracy and rights, and a musical group reaching speakers through song.
  • Universities typically contribute methodological rigour, ethics infrastructure, linguistic expertise, trainable students and institutional continuity. Community organisations contribute access, trust, variety knowledge, appropriateness judgement and validation against lived experience.
  • FAIR and CARE principles, alongside PILARS, are named as the standards for long-term sustainability in national language data commons infrastructure.
  • Indigenous data sovereignty frames data as self-determination, with communities seeking authority over collection, use and sharing, and defines data to include culture, language, land and relationships rather than only statistics.
  • Recommended focus areas include community and project context, the changing digital landscape, individual and collective knowledge protections, planned outputs, and confidentiality nuances.
  • Individual consent does not settle collective knowledge rights, which is the distinction commercial consent frameworks handle worst.
  • MILPA has run since 2015 between UC Santa Barbara and diasporic Mexican Indigenous community organisations.
  • CoEDL and AIATSIS partnered from the outset with a named liaison role, and reciprocity in both directions, improving research access while helping community groups deposit and safeguard material.
  • The New Commons Indigenous Language Data Commons Incubator, co-designed with a 19-member global Steering Committee and implemented with the GovLab and Microsoft, supports Indigenous-led teams building communitygoverned datasets.
  • Partnerships take longer to establish than projects take to run, so relationships should be built ahead of demand rather than in response to a contract.

Sources and further reading

Frequently asked questions

Because they contribute different things.

Before collection. The Puno Quechua project set its CC0-1.0 licence during preparation, ahead of any data existing. Deciding afterwards means renegotiating with everyone who contributed.

FAIR covers findability, accessibility, interoperability and reusability. CARE covers collective benefit, authority to control, responsibility and ethics, and exists because FAIR principles alone can facilitate extraction.

Individual consent covers a person's own contribution. A song, story or ceremonial term may belong to a community rather than to the individual who recorded it, and individual consent does not settle that.

Longer than a project. MILPA has run since 2015 and CoEDL partnered with AIATSIS from the outset of a multi-year centre. Relationships should be built ahead of demand rather than in response to a signed contract.

Yes, provided the commercial position is stated plainly.

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team