Battlefield Data for AI: Ethics, Palantir, and Surveillance
Ukraine’s drone war is generating a body of machine-readable experience of extraordinary value. Concretely, 500,000 hours of drone-collected footage for machine learning. The urgent question for the data industry is not whether that experience can improve civilian AI, but under what authority (and with what safeguards) it should ever be allowed to do so.
War and battlefield data as a dataset where humans train software
By Manuel Herranz · Founder and CEO, Pangeanic
The moral supply chain behind battlefield AI
Every war is a brutal school. Armies have always studied what worked, what failed and what the enemy did next. They have turned experience into doctrine, redesigned equipment and carried military inventions into civilian life. Radar, satellite navigation and even the internet all have such genealogies. There is nothing new in war producing knowledge. What is new is the form in which that knowledge is now produced: not only as reports written after an operation, but as continuous, synchronized and increasingly annotated streams of video, thermal imagery, telemetry, coordinates, operator inputs and outcomes.
The battlefield is becoming a data factory
That is the unsettling achievement described by MIT Technology Review in its report on the emerging market for Ukrainian drone data.[1] Tens of thousands of missions can generate millions of observations of machines and people acting under conditions that are difficult, dangerous or impossible to reproduce in a laboratory or under an RFQ. It’s more than “real life”, and for an AI developer, this is exceptionally valuable material. It contains occlusion, interference, low visibility, damaged hardware, improvised tactics, and adversarial behavior: the long tail of reality for which conventional datasets are usually impoverished.
Its value comes from more than authenticity. Battlefield collection produces edge cases at scale, combines visual, thermal and operational signals, records decisions made under pressure and compresses time: failure modes that a civilian test program might encounter over years can accumulate in months. War is, among its many obscenities, an accelerator of variance. That density makes the data attractive to engineers precisely because it makes the ethical questions difficult to postpone.
It is also a record of human beings living, fighting, fleeing, rescuing, and dying. I’m not being sentimental, but we should determine what the “data-for-AI” industry is permitted to do next. And in line with the ethical side of AI.
Ukraine never chose to become an experimental laboratory. It is defending itself against invasion, and the rapid development of AI-assisted systems may confer legitimate operational advantages, protect soldiers, and improve the precision of some military decisions. Any ethical analysis that ignores that necessity becomes an exercise in comfortable casuistry. The primary responsibility for the war lies with the aggressor, not with the society forced to innovate under fire.
But military necessity at the point of collection is not a perpetual commercial license. A lawful purpose in one context does not automatically authorize every reuse in another. If battlefield data migrates into delivery drones, agricultural monitoring, border systems, urban cameras, or police robotics, the original emergency has crossed into a different social and legal order. The important question is no longer only, “Was this data useful in war?” It is, “Who may use it after war, for what purpose, and who remains accountable when the original context has disappeared?”
Avengers Labs is a separate Ministry platform, developed by the team behind the DELTA battlefield-management system. Its core is an annotated collection of five million battlefield frames, most drawn from DELTA and continuously replenished, covering equipment, infantry and aerial targets. Under Cabinet Resolution No. 310, approved participants can obtain licenses to train, configure, test and evaluate models without receiving direct extractive access to the underlying Ministry databases. Agreements can specify territory, duration, monitoring, liability and subcontracting, while transfers of resulting military-purpose products remain subject to Ukrainian export control. The first Ukrainian companies signed licenses in August; the United Kingdom then became the first foreign government partner.[4][21]
There are very defensible reasons to prefer a controlled data room to the indiscriminate export of raw footage. Bringing an algorithm to the data, logging access and keeping sensitive material inside a protected environment can reduce leakage. Indeed, the EU AI Act itself recognizes privacy-preserving architectures in which algorithms are brought to data rather than data being copied between parties.[5] But a secure room solves only one part of the problem. Once a model has learned from the material, the capability (and sometimes information about the training data) can leave in the weights, embeddings, thresholds and derived products. The raw file may remain in Kyiv while its statistical palimpsest travels globally.
Nor does permission to enter the room settle who controls what comes out. A license must address fine-tuned weights, adapters, embeddings, distilled or merged models and synthetic data generated from them, not only the source files. Otherwise, provenance can be laundered one generation later: a conflict-trained model produces an apparently synthetic corpus whose origin no longer appears in the dataset card. Data sovereignty that stops at the database door is incomplete.
Operational security travels in the same direction. A derived model or detailed evaluation can reveal which signatures are detectable, where sensors fail, what confidence threshold triggers action, and which countermeasures work. A model can be both a capability and a compressed intelligence leak. Export review must therefore examine what the derivative discloses, not merely whether a raw video file crossed a border.
This is why the data industry must stop treating model training as a moral washing machine. Transformation does not extinguish provenance. A video does not become ethically immaculate because it has been converted into vectors, nor does a questionable source become acceptable after enough annotation invoices have changed hands.
The commercial market makes this especially urgent. Enabled Intelligence, a US data-labeling company, advertises access to more than 500,000 hours of pre-labeled electro-optical and infrared footage from the war in Ukraine. Speaking as the representative of a company that provides ethical data for AI development… What is the provenance of the dataset? How the hell did they get half a million hours when most RFQs and requests for data only request a few thousand?
Enabled Intelligence’s catalog describes ontologies for aerial objects, vehicles and ground activity; the company has also pointed to applications beyond defense, including remote sensing and commercial delivery. When DefenseScoop asked where the footage came from, the company did not disclose the source, its government customers, or the agencies it supports.[6]
We’re not talking about a marginal data broker operating at the edge of the defense economy. In 2025, the US National Geospatial-Intelligence Agency awarded Enabled Intelligence the single-award SEQUOIA contract, with a ceiling of $708 million over up to seven years, for data labeling supporting geospatial AI across defense and intelligence programs.[26] A contractual ceiling is not guaranteed revenue, but the scale of the vehicle reveals how central annotation and dataset preparation have become to contemporary intelligence infrastructure. The person drawing the bounding box may now sit surprisingly close to the strategic center of gravity.
That asymmetry is extraordinary. A buyer can apparently learn the annotation format and request a price, but the public cannot learn the source, chain of title, collection authority, or downstream restrictions attached to the underlying human record. This is technically mature but institutionally juvenile: a familiar imbalance in the AI economy, only here the stakes are measured in lives rather than clicks.
It would be irresponsible to leap from this market’s existence to the allegation that a particular company wants the war to continue; no such claim can be credibly supported. However, no such motive needs to exist for an extractive structure to emerge. The narrower incentive problem is real enough. If compensation is tied to hours, frames, or the novelty of combat conditions, commercial value rises with the volume and distinctiveness of records produced by violence. Contract design should not reward the generation of more battlefield material. Benefit sharing can instead be tied to downstream revenue, Ukrainian technical capacity, and reconstruction.
Palantir’s involvement is a matter of public record. Ukraine states that Brave1 Dataroom runs on Palantir software. President Volodymyr Zelensky met Palantir CEO Alex Karp in May 2026 as Ukraine expanded cooperation in battlefield analysis, intelligence processing and deep-strike planning; Reuters reported that more than 100 companies were then training over 80 models through Brave1.[7] Palantir also has a documented, separate partnership with Enabled Intelligence: users of Palantir Foundry in US government environments can send datasets into an integrated Enabled Intelligence labeling workflow and receive the annotated data back inside Foundry.[8]
The operational depth is not merely inferred from architecture diagrams. In February 2023, Karp said Palantir was “responsible for most of the targeting in Ukraine,” while a company spokesperson described its software as helping Ukrainian forces target tanks and artillery.[27] By May 2026, Karp was more precise: Palantir formed part of Ukraine’s targeting system, he said, but much of that system had been built by Ukrainians themselves.[28] Then, on September 1, Reuters reported that former Ukrainian defense minister Mykhailo Fedorov had named Karp as the first major investor in a new Ukrainian defense-technology company.[29] None of this proves ownership of the footage at issue, but it demonstrates a deep relationship, even at a personal level, that now spans software, operations, model training and capital.
President Zelenskyy fired Mr Federov (himself, both an entrepreneur and a politician) in the summer of 2026.
These facts establish a substantial role for Palantir in the wider ecosystem. They do not establish that Enabled Intelligence obtained its Ukrainian footage through Brave1, that Palantir owns the collection, or that Palantir has reused this particular material in civilian products. No public evidence I found completes that chain. Yet the dataset is there, ready to be purchased if you have a deep enough pocket.
The undisclosed source is therefore not permission to invent an answer; it is evidence of a transparency failure that the supplier and its customers should remedy.
Musk's technologies call for a different conclusion. SpaceX’s Starlink is a critical part of Ukraine’s communications infrastructure. Ukraine has relied on tens of thousands of terminals for battlefield communications and, in some cases, for piloting drones.[9] Reuters has also reported disputes over SpaceX’s ability to restrict service, an account the company contested.[10] This creates a serious strategic dependency question: a private infrastructure provider may possess considerable operational leverage over a sovereign state at war.
It does not, however, prove ownership of the data moving across that infrastructure. The fact that a network transports a packet does not establish title to its payload. I found no credible public evidence that SpaceX, xAI, or Tesla operates the marketplace described by MIT Technology Review, supplied the 500,000-hour collection, or holds rights to reuse it. As of this writing, the responsible description is as follows: Starlink serves as the communications layer; Palantir is a confirmed software and analytics layer; Enabled Intelligence is a data-conditioning and distribution actor. Those roles may overlap in practice, but they should not be conflated without evidence.
The commercial-military loop and the migration of purpose
The most important risk is not that every military dataset is inherently illegitimate. It is that data gathered under one set of norms can migrate into systems governed by another. Privacy scholar Helen Nissenbaum describes privacy not simply as secrecy but as the appropriate flow of information within a social context.[11] On that account, a transfer can violate privacy even when the information was lawfully visible or collected at the outset. Battlefield-to-city reuse is an extreme case of this collapse of context.
Consider a thermal model trained to detect a concealed person beneath foliage. In one setting, it could help locate a missing child or an earthquake survivor. In another, it could help a border authority track migrants, an authoritarian government identify dissidents, or a police force conduct persistent surveillance without individualized suspicion. The underlying capability is similar; the purpose, power relationship, safeguards, and consequences are not.
The transfer has already begun at the level of products and practices. Drone manufacturers have used operational experience in Ukraine to improve systems subsequently marketed to civilian public-safety agencies. WIRED reported, for example, that BRINC drew lessons from Ukraine while developing a drone that was later adopted by US police departments.[12] In the other direction, Ukraine began using Clearview AI’s civilian facial-recognition technology during the war for functions including checkpoint screening and identifying the dead.[13] These examples do not show that every transfer is abusive. They show that the wall between military and civilian technical ecosystems is highly permeable.
The result is not a simple one-way transfer from military research to civilian prosperity. Civilian technologies are adapted for combat; combat generates unusually dense operational data; commercial firms structure that data; and the resulting capabilities return to civilian markets. The loop may improve navigation, resilience and safety, but it can also carry military assumptions about threats, acceptable error and human oversight into environments where the state’s powers (the citizen’s rights) are supposed to be different.
The crossing point is already visible
The August UK–Ukraine agreement makes that permeability unusually concrete. British companies Sintela, Mind Foundry and Skyral are piloting a system that uses fiber-optic cable as an AI-enabled sensing array, developed with access to Ukrainian battlefield data. Its first announced deployment is at a British defense facility, while UK officials have identified airports, prisons, railways and energy infrastructure as possible later settings.[22]
This does not show that a target-recognition model has been lifted unchanged from a drone and installed beside a prison, or that all military and civilian security uses are morally equivalent. It shows something more (administratively) important: a procurement path now exists along which conflict-derived data, learned capabilities and evaluation practices can move into environments with different people, laws and acceptable error rates. At each crossing, authorization and testing should be renewed. Provenance must travel with the capability, not only with the raw files.
Surveillance risk also resides in the labels, not only in the pixels. “Person,” “vehicle,” “civilian,” “hostile,” “target”, and “anomalous behavior” are not neutral descriptions waiting to be discovered in nature. They are institutional judgments made under uncertainty, sometimes under extreme time pressure and within a particular doctrine. Annotation converts those judgments into an ontology that a model can reproduce at scale. When the model crosses into civilian life, it can import the battlefield’s prior assumptions about who or what deserves attention. You can expect to become "hostile" because you walked through a zebra crossing if the neural net (in itself a black box and now with military input into the model) makes the necessary calculations.
This is both an ethical problem and a data-quality problem. Combat footage may be rich in rare events, but it is not automatically representative of peaceful environments. It is shaped by mission selection, geography, weather, sensors, tactics, operator behavior, and survivorship. It is also generated in the presence of an opponent actively trying to corrupt the observer’s picture. Decoys, camouflage, spoofed signals, electronic warfare, and selective recording are not noise around the dataset; they are part of its generating process. Errors may be hard to identify because the “ground truth” of a strike or classification is itself contested. Conflict labels should therefore record confidence, observer, verification method, and disputed status, rather than converting uncertainty into a clean class ID. Real-world data is not the same thing as unbiased data. A system trained on an adversarial theater may perform impressively on a benchmark and still behave dangerously when applied to traffic management, emergency response, or public surveillance.
This is why serious dataset documentation must describe motivation, composition, collection, labeling, intended uses, and excluded uses, the logic advanced by the “Datasheets for Datasets” research tradition.[14] For conflict-origin data, ordinary documentation is necessary but insufficient. The source conditions are so coercive, the downstream capabilities so dual-use, and the affected people so unable to object that a higher standard is warranted.
Preserve the evidence before optimizing the dataset
Some battlefield footage may have three incompatible futures: operational intelligence, AI training material, and evidence of possible violations or civilian harm. A training pipeline optimizes for utility. It clips, crops, compresses, blurs, normalizes, relabels, removes duplicates, and often strips metadata. An evidentiary workflow optimizes for authenticity and context. It preserves the original bits, timestamps, location and device metadata, cryptographic hashes, access history, and chain of custody. If the training workflow comes first, some of that value can be destroyed irreversibly.
The discipline set out in the OHCHR–UC Berkeley Protocol for handling digital information in investigations points toward a preservation-first rule.[23] Before human-bearing conflict data enters a commercial training pipeline, a competent authority should assess whether it has evidentiary or humanitarian value, retain an immutable original under appropriate access controls, and create a separate, traceable training derivative. The same caution applies to footage and metadata that may help identify the dead or missing. The ICRC describes protection of personal data in armed conflict as part of safeguarding life, integrity, and dignity; tactical usefulness does not exhaust the human interests in the record.[24]
Consent is not the whole answer
The language of consent is attractive because it appears to provide a clean ethical switch: consent obtained, use permitted; consent absent, use prohibited. War exposes the limits of that assumption. A civilian captured incidentally by a reconnaissance system cannot meaningfully negotiate with the aircraft overhead. A soldier may be acting under military authority. A person classified as a target will plainly not consent. Even where individual consent is impossible, some military processing may still be lawful under international humanitarian law and domestic authority.
The relevant test must therefore be broader: lawful authority, military necessity, distinction and proportionality at collection; then purpose limitation, minimization, security, accountability and human-rights due diligence for retention and reuse. The absence of consent does not make all wartime collection unlawful. Nor does the presence of state authorization make all subsequent commercialization ethical.
European law illustrates the discontinuity. The EU AI Act excludes systems developed or used exclusively for military, defense or national-security purposes. Yet it expressly brings systems back within scope when they are used for civilian, humanitarian, law-enforcement, or public-security purposes.[15] For high-risk civilian AI, Article 10 requires governance over the origin of training data, the original purpose of personal data collection, annotation processes, assumptions, bias, and data gaps.[5] The GDPR separately requires lawfulness, fairness, transparency, purpose limitation, minimization, security, and accountability when personal data falls within its scope.[16] The European Data Protection Board has warned that personal data may be absorbed into model parameters and that an AI model is not automatically anonymous merely because the original records are no longer directly displayed.[17]
Europe also has an underused disclosure lever. Article 53(1)(d) of the AI Act requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content, using a template issued by the European Commission.[25] The current template is not a conflict-data register, and disclosure cannot manufacture consent or lawful authority. But a declared category for operational data derived from armed conflict (covering the authorizing state, broad data types, permitted purposes, and restrictions)could make some downstream crossings visible without creating a new regulator. The disclosure should be tiered: a categorical public summary, a confidential annex for a regulator or accredited auditor, and continued protection for classified operational detail. Radical transparency would be irresponsible in wartime; permanent opacity is not the only alternative.
International humanitarian law, data-protection law, export controls, procurement rules and the AI Act therefore form a patchwork, not a complete vacuum. The problem is that no single regime reliably follows conflict-origin data through collection, annotation, model training, derivative models, sale, and eventual civilian deployment across jurisdictions. Military exceptions and commercial secrecy create seams. Data supply chains are exceptionally skilled at finding seams.
The ICRC has called for limits on autonomous weapons because of the risks posed by unpredictability and the loss of human control, while NATO’s own responsible-AI principles include lawfulness, accountability, explainability, traceability, reliability, governability, and bias mitigation (beyond the traditional 4 pillars of Ethical AI).[18] These principles should not stop at the deployed weapon. They should travel backward into the dataset and forward into every material derivative.
We need a conflict-data protocol
There is a useful, if imperfect, precedent in the governance of conflict minerals. Of course, minerals and data are not morally or economically identical. But when a valuable input originates in a conflict-affected environment, ordinary supplier assurances are inadequate. The OECD’s framework asks companies to build management systems and traceability, identify and mitigate risks, submit to independent audits, and report publicly on their due diligence.[19]
The data industry needs an equivalent discipline for conflict-origin training material. I would propose 7 obligations.
Persistent provenance, evidentiary preservation, and chain of authority. Every conflict-origin dataset should carry a verifiable record of who collected it, under what authority, through which transfers, with which transformations, and under whose control. An immutable evidentiary original should be preserved before a training derivative is created. “Source confidential” may occasionally protect operations; it cannot serve as a permanent substitute for due diligence by purchasers, auditors and regulators. The record must also distinguish state authorization, intellectual-property rights, the interests of people depicted and custody of potential evidence; “ownership” is rarely a single checkbox.
Purpose-bound authorization. Licenses must identify permitted objectives and prohibited domains. A defense authorization should not silently expand into permission for mass surveillance, facial or gait recognition, migration enforcement, employee monitoring, insurance decisions or unrestricted resale. Any transition to a civilian purpose should trigger a fresh legal basis, necessity assessment, and fundamental-rights review.
A risk taxonomy and a test of necessity. Environmental telemetry and sensor performance information can often be separated from imagery used to identify or classify people. Human-bearing video, faces, voices, gait, precise locations, casualty imagery, and disputed target labels should be treated as presumptively noncommercial. De-identification must be tested, not declared, and the least rights-intrusive data capable of achieving the legitimate objective should be used. Before ingesting human-bearing battlefield material, a developer should also demonstrate why simulation, synthetic data, or nonhuman telemetry cannot provide a sufficient alternative. Synthetic data is not ethically automatic—it can reproduce the assumptions built into its generator, but it can reduce unnecessary exposure of real people to repeated downstream use.
Secure computation plus derivative controls. Data rooms, clean rooms, and access logs are good architecture, but they must be paired with evaluation of memorization, extraction, poisoning, and model-exfiltration risks. Providers should maintain a registry of which models and versions learned from which conflict-origin sources. Contractual controls must remain attached when a model is licensed, merged, fine-tuned, or distilled and when it generates synthetic training data; they should define export, sublicensing, retention, audit, revocation, and the practical limits of deletion or retraining.
Independent review and tiered disclosure. High-risk conflict datasets should undergo legal, technical, and human-rights review by bodies independent of the commercial transaction. Public reporting can protect operational details while still disclosing source categories, collection authority, affected populations, permitted uses, audits, incidents, and downstream product classes. More sensitive facts can be disclosed confidentially to regulators or accredited auditors. A company selling precision annotations while pleading imprecision about provenance has its priorities inverted.
Sovereign control and benefit sharing. Ukraine should not become merely the quarry from which foreign companies extract digital value. The state must retain meaningful control over access and derivative use, subject to human rights safeguards, and the economic value generated by its wartime experience should support Ukrainian capacity, security, and reconstruction. Sovereignty is not satisfied because the raw archive remains on Ukrainian servers; it also concerns who can operate, update, suspend, export, or revoke access to the resulting models. Benefit sharing does not cure an otherwise impermissible use, but extraction without it repeats an old political economy in a new medium.
Refusal, revocation, and remedy. Ethical sourcing requires the capacity to say no. Contracts should permit access to be suspended, datasets withdrawn, and affected models retrained or retired when provenance claims fail, risks change, or restrictions are breached. Machine unlearning remains difficult and imperfect; that is an argument for designing revocability before ingestion, not for declaring withdrawal impossible after the fact. There must be an incident process, an accountable decision-maker, and a route for affected people or their representatives to raise grievances. A supplier that cannot reject a dataset is not exercising governance; it is merely operating a pipe.
These duties should not fall only on the original collector. Annotators, data brokers, cloud and platform providers, model developers, systems integrators, procuring agencies, and civilian deployers each occupy a position of control. Responsibility should be proportionate to knowledge, leverage, and capacity to prevent harm. Palantir itself has argued that technology providers bear greater responsibility as they become more deeply embedded in military operations.[20] That proposition is sound; it should also apply to downstream data reuse.
Where should an ethical data provider draw the line?
After more than 26 years of building data and language technology, I have learned that “data quality” cannot be reduced to accuracy, volume, and format. A technically pristine dataset can still be ethically defective. Quality also includes provenance, lawful authority, contextual integrity, representativeness, traceability, and the terms under which others may use the result.
For us at Pangeanic, the conflict-origin would therefore be a first-class governance attribute, not a footnote in a procurement form. We would require evidence of lawful sourcing and chain of title; classify personal, biometric, and human-bearing content; document annotation decisions and uncertainties; impose purpose-specific downstream restrictions; preserve lineage in derived models; and require secure processing, human oversight, and audit. Where those conditions could not be established (or where the risk of civilian surveillance could not be credibly contained), the correct commercial decision would be to refuse the data.
This position does not deny that battlefield learning can produce civilian benefit. Better navigation in degraded environments can aid disaster response. Robust perception can make industrial inspection safer. Improved remote sensing can support agriculture and environmental protection. The mistake is to assume that a beneficial possibility cancels provenance. It does not. Dual use is a reason for stronger governance, not a euphemism that ends the discussion.
Silicon Valley has spent years telling us that data is the new oil and tokens are the new coal. War has now supplied a refinery, although the crude contains the recorded experience of people who never agreed to become a product. The metaphor should make us uncomfortable because the market logic behind it is already familiar: acquire first, normalize later, and let opacity perform the work that consent cannot.
War may teach machines. It must not be allowed to rewrite the civilian social contract by default. The data industry’s test is not whether it can make wartime data useful. It is whether it can prevent necessity at the front from becoming normality at home.
[23] Office of the United Nations High Commissioner for Human Rights and Human Rights Center, UC Berkeley School of Law, Berkeley Protocol on Digital Open Source Investigations, 2022. The Protocol concerns digital open-source investigations; this essay applies its preservation logic by analogy to battlefield sensor records that may later have evidentiary value.