Home Guide
How to license your company’s data for AI training
What AI labs want from operating companies, which data qualifies, and the steps from a first call to a signed license.
Last updated October 2026 · 7-minute read
Short answer. A company licenses its data for AI training by granting an AI lab, or a data company that supplies labs, the right to use a defined copy of its internal work data: chat, tickets, documents, code reviews and call recordings. The company keeps ownership and keeps using the data. The license sets which sources go in, what is de-identified, how long the buyer may keep raw data and when the company is paid.
Why AI labs want company work data now
Language models already know the public web. What they lack is a record of how real teams work: a support engineer reproducing a bug, a product manager turning feedback into a ticket, a developer answering review comments, a team deciding what ships on Friday. AI labs are now training models to act as agents inside companies, and that work is learned from years of everyday operational data that only companies hold.
That is why the data companies that supply AI labs have started to license internal archives from ordinary businesses, not only from publishers and platforms. Software and services companies are a natural fit: their work is written down, in English, in tools with clean exports.
Selling your data or licensing it
People search for how to “sell” company data, but the deals that make sense for an operating company are licenses. In a license:
- the company keeps ownership and keeps using its data;
- the buyer gets the right to use a prepared copy, for stated purposes such as training and evaluation;
- the scope is written down source by source, and anything not listed stays out;
- the license can be non-exclusive, so the same data can be licensed again, or exclusive for a limited time and scope.
An outright sale usually only happens when a company winds down and its archive is sold as an asset. For a company that is still operating, a license is the safer and more flexible route.
Which companies qualify
Buyers look for depth and cleanliness more than size. A company is usually a good fit when it has:
- 30 or more people, now or at its peak. Thirty to 200 is the sweet spot for most buyers.
- Three or more years of history in its work tools.
- Mostly English content. Buyers check each source, and many need about 90% English.
- The rights to license the data: its own internal work first, and client project work only where contracts allow it.
- A team that has been told. Reputable buyers ask for it, and it avoids surprises later.
US-based companies are preferred by many buyers; companies in the UK, the EU and Canada are often considered too. The 2-minute self-check scores these points for your company.
Which data qualifies, and what stays out
The most useful sources are the ones that show work being done and decisions being made:
- Team chat (Slack, Microsoft Teams): working threads, not HR or personal channels.
- Work tracking (Jira, Linear, Asana): tickets with history, acceptance criteria and links to code.
- Code (GitHub, GitLab): code you own and its review history.
- Documents (Confluence, Notion, Google Drive): specs, runbooks, postmortems, plans.
- Support and sales (Zendesk, HubSpot): tickets and conversations, heavily de-identified.
- Call recordings, only with consent and usually with the strictest handling.
Health, payment and personal finance data usually stays out. So does client code or client data that your contracts don’t let you license, open-source code you didn’t write, and anything covered by a third party’s confidentiality you can’t clear.
The process, step by step
- A first check. A short call on your tools, history, language and rights tells you whether there is a deal worth doing.
- An NDA. A mutual NDA is signed before you share any detail about your data.
- Scope and rights. You decide source by source what goes in and what stays out, and confirm you have the right to license each one.
- The buyer that fits. Different buyers want different data, volumes and company locations. You get introduced to the one that fits yours.
- The license. You sign directly with the buyer. Your counsel reviews it; the terms that matter most are below.
- Export and review. The buyer’s team exports and de-identifies the data as the license requires, reviews what was delivered, and pays. This often takes 60 to 90 days after signing.
Three ways to do it
| Route | How it works | Trade-off |
|---|---|---|
| Listing on a data marketplace | You list a dataset and wait for buyers. | Little help on scope, rights or terms; suits ready-made datasets more than work archives. |
| Going directly to a buyer | You approach a lab or data company yourself. | You negotiate alone, without knowing what similar datasets go for or which terms are normal. |
| An advisory on your side | An adviser checks the data, scopes it, introduces the buyer that fits and stays through the contract. | You share the deal with an adviser; with a buyer-paid model, it costs the company nothing. |
What to check before you sign
Six terms decide most of the value of a data license: the liability cap and whether indemnities sit inside it; the consents and rights you warrant; exclusivity, both how long and over what; acceptance criteria and the review window; what happens if part of the data is rejected; and when you get paid. The FAQ goes through each, and your own counsel should sign off on all six.
How Gemdata helps
Gemdata Advisory is an independent, seller-side advisory. We tell you on a 15-minute call whether your data qualifies, describe it the way buyers read it, introduce you to the buyer that fits and help you on price and terms until you are paid. We advise on calls only and never touch your systems or data. It is free for the company: when the deal closes, the buyer pays us a finder’s fee.
This guide is general information, not legal or tax advice.