For years, training data for large languagemodels had a fairly clear job: help models understand language, acquireknowledge, and follow instructions. Whether the data came in the form ofquestion-and-answer pairs, instruction data, or preference data, it generallyrevolved around the same basic idea: given this input, what should the modelsay or do in response?
That assumption is starting to change.
As AI moves beyond conversation and intoreal-world task execution, models are expected to do much more than generate agood answer. An AI agent may need to understand a user’s goal, break it intomanageable steps, choose the right tool, call it with the right parameters,interpret what comes back, adjust its next move, recover when something goeswrong, and ultimately deliver a result that can be checked and verified.
In other words, the model is no longer justlearning how to respond. It is learning how to act.
The National Data Administration’sImplementation Plan for Advancing the Development of High-Quality IndustryDatasets—released in June 2026—explicitly calls for faster development ofdatasets for complex task planning, long-horizon reasoning, human-computerinteraction, and decision-making and execution, supporting emerging AIapplications such as AI agents.
From Glodom’s experience across languageservices and AI data projects, this shift goes much deeper than adding a fewnew fields to existing datasets. The underlying logic of training data ischanging: models have traditionally learned how to answer; increasingly, theyalso need to learn how to get something done.
And to teach a model how to get somethingdone, the data needs to capture more than the final answer. It needs to capturethe actions that led there.
1. From Answers to Task Trajectories
In conventional language model training, asingle data sample might be as simple as a user question paired with a modelanswer.
That is rarely enough for an AI agent.
Consider a customer service request such aschecking an order and changing the shipping address. A useful training examplemay need to show how the agent recognizes the customer’s intent, which tool itchooses first, what parameters it sends, how it interprets the tool’s response,what it does when information is missing, and how it recovers if a tool callfails.
The value lies in the sequence, not justthe endpoint.
This is where the concept of a tasktrajectory becomes important. A trajectory can capture the entire chain ofactions involved in completing a task—from understanding the goal and planningthe workflow to using tools, interpreting feedback, deciding what to do next,and reaching the final outcome.
Every part of that process may have valuefor model training, fine-tuning, or evaluation.
The National Data Administration’s recentpolicy guidance reflects the same shift. Alongside traditional multimodal datasuch as text, code, images, audio, and video, it highlights the need to developdatasets for complex task planning, long-horizon reasoning, human-computerinteraction, and decision-making and execution.
For AI data providers, this changes the jobitself.
The question is no longer simply “What isthis piece of content?” It is also “Why was this action taken?” and “Whatshould happen next?”
That difference is one of the clearest waysin which AI agent data is beginning to move beyond traditional annotation.
2. The Hard Part Is Defining What a “Correct” Action LooksLike
The challenge with AI agent data is notsimply that the samples are longer.
The bigger issue is that quality has becomeharder to define.
Traditional text annotation often focuseson things such as semantic accuracy, intent, or whether an answer is correct.In an agentic workflow, however, a trajectory does not become high-qualitysimply because the task was eventually completed.
An agent might choose the wrong tool andrecover later. It might take several unnecessary steps before reaching the sameresult. In more sensitive cases, the final outcome might look correct eventhough the process itself has violated a business rule or crossed a securityboundary.
That means agent data needs to be evaluatedfrom several angles at once: Was the task completed correctly? Was the pathappropriate? Were the actions efficient? Did the agent stay within the requiredsecurity and permission boundaries?
And those questions cannot be answered inisolation from the business context.
A financial AI agent handling an accountoperation needs to understand business rules and permission levels. Ahealthcare agent needs to recognize specialized terminology and distinguishbetween actions with very different levels of risk. An industrial agentinteracting with equipment has to work within real operating procedures andsafety requirements.
This is why Glodom places such emphasis onbringing together language expertise, technical capabilities, and industryknowledge in AI data services.
With more than 20 years of experience inlanguage services, Glodom has built language resources covering more than 200languages and 40 countries and regions, alongside project experience across awide range of industries. As AI applications have evolved, we have extendedthat expertise into AI data services, combining linguistic capabilities withdata collection, cleaning, processing, annotation, dataset development, andquality management.
For AI agents, this combination mattersbecause many of the problems they face sit precisely at the intersection oflanguage, tools, and business processes.
3. Human-AI Collaboration Is Changing How Data GetsProduced
The data production process itself isbecoming more intelligent.
The National Data Administration reportedthat, as of the first quarter of 2026, more than 116,000 high-quality datasetshad been developed nationwide, with a total volume exceeding 960 PB. Averagedaily token consumption had also surpassed 140 trillion. As the amount of datacontinues to grow, producing everything through conventional manual workflowsalone makes it increasingly difficult to balance speed, scale, and quality.
The answer is not simply to “use more AI.”
The National Data Administration’simplementation plan calls for data annotation to evolve from a primarilyhuman-led approach toward human-AI collaboration with deeper expertinvolvement. It specifically highlights models such as “model pre-annotation +human calibration” and “human annotation + model verification.”
In practice, the real value of human-AIcollaboration is less about eliminating manual work and more about decidingwhich work should be done by machines and which work still requires people.
Highly standardized tasks—format checks,repetitive processing, and preliminary classification, for example—can often beaccelerated through models and automation. But when a sample requires nuancedsemantic judgment, an understanding of business rules, careful assessment oftool-use paths, or decisions around security boundaries, experiencedprofessionals still play a critical role.
Glodom takes a similar approach in its AIdata services. Technology helps us increase processing efficiency, whileprofessional teams focus their attention on critical samples that requiredeeper review and calibration. The goal is not to choose between automation andhuman expertise, but to use each where it works best.
The broader policy direction points to thesame trend. As the National Data Administration calls for more specialized andintelligent annotation, with experts playing a deeper role in data productionfor instruction fine-tuning and reinforcement learning, data work is becomingincreasingly knowledge-intensive and technology-intensive.
4. Multilingual AI Agents Add Another Layer of Complexity
For AI products built for global users,there is another challenge that cannot be ignored: language differences do notdisappear simply because everyone is using the same underlying model.
A customer service task may look completelydifferent from one market to another. People describe orders, refunds, oraddress changes in different ways. Even the same phrase can carry differentimplications depending on the local business context and may therefore lead toa different workflow.
That makes multilingual agent data muchmore than a translation exercise.
Translating a Chinese dataset into English,Japanese, or German does not automatically produce a useful multilingualtraining set. The wording has to be accurate, but so does the underlyingintent. Business workflows need to map correctly, and tool-use logic needs toremain consistent with the task being performed.
There is also the question of datasecurity.
Real-world business datasets can containpersonal information, transaction details, and internal operational data.Requirements around collection, cleaning, anonymization, storage, and data usetherefore need to be considered throughout the entire workflow.
This is an area where experience inlanguage services can translate directly into value for AI data production.
Through years of working with globalclients, Glodom has developed experience in multilingual data collection andprocessing, together with capabilities spanning text, speech, image, and videodata. Depending on project requirements, we can support the full workflow fromdata collection and cleaning to OCR/ASR, structured processing, annotation, andquality review. Projects can also be delivered in line with informationsecurity requirements and private deployment needs.
For multilingual AI applications, languageis not an extra layer added after the data has been produced.
Language is part of the data itself—andlanguage quality can ultimately affect how well an AI system performs in thereal world.
5. From Producing Data to Producing Data That Can ActuallyBe Used
The AI data services industry is entering adifferent stage of development.
The question used to be relatively simple:Do we have enough data?
Now there is a second question that mattersjust as much: Is this data actually useful for training models and supportingreal applications?
The National Data Administration has set agoal of building, by the end of 2028, a number of high-quality industrydatasets covering key areas and validated through real-world applications. Italso aims to create a continuous cycle in which scenarios drive data, datadrives models, models support applications, and applications generate value.
That is why AI data services can no longerbe fully described by terms such as “data collection” or “data annotation”alone.
An AI agent operating in a real businessenvironment needs data that reflects different ways users express themselves,different task paths, tool calls, exceptional situations, and business rules.For global AI applications, those datasets also need to account for differencesacross languages and cultural contexts.
At Glodom, we see this as one of thecentral challenges for the next stage of AI data services.
Building on more than two decades ofexperience in language services and our growing capabilities in AI data, we aimto deliver more than processed datasets. The goal is to provide data assetsthat can genuinely support model training, evaluation, and real-world AIapplications.
As AI moves from talking to doing, trainingdata has to evolve with it—from capturing an answer to capturing the fullcourse of a task.
Because, in the end, what an AI modellearns depends to a great extent on the kind of experience its data allows itto learn from.

