Open data is information that anyone can freely access, use, modify and share for any purpose, subject at most to a duty to credit the source and to keep it open. The Open Knowledge Foundation's Open Definition, now at version 2.1, sets that test. Three conditions follow from it: an open licence that grants reuse rights in advance, a machine-readable format that software can process without retyping, and access with no more than a marginal reproduction fee. The idea exists because data collected at public expense, or pooled by volunteers, is worth more when anyone can build on it than when it sits behind contracts and price lists.
Most open data comes from governments: maps, weather observations, company registers, census tables, transport timetables. Volunteer projects such as OpenStreetMap and Wikipedia, and a growing set of corporate releases, add to the pool. Data that is merely visible on a website is not open data. Without a licence, reuse remains legally uncertain however easy the scraping.
How data becomes open
The licence does the legal work. CC0, released by Creative Commons in 2009, waives all rights and places a dataset in the public domain. CC BY 4.0 requires only attribution. The Open Database License (ODbL), published by Open Data Commons on June 29, 2009, adds share-alike: anyone who publishes a derived database must release it on the same terms. Governments often write their own. The UK's Open Government Licence (OGL), maintained by The National Archives, reached version 3.0 on October 31, 2014 and is interoperable with CC BY 4.0. In the United States, works produced by federal employees carry no copyright at all.
Format determines usability. A spreadsheet in CSV, a JSON feed or a GeoJSON map layer can be loaded directly into analytics tools; a scanned PDF technically published "openly" cannot. Publishers increasingly offer two routes. Bulk download delivers the whole dataset as files. An application programming interface (API) returns specific records on request and suits data that changes by the minute. Portals such as data.gov in the US and data.europa.eu in the EU catalogue the files.
A worked example shows the chain. Since February 2, 2026, every UK petrol station must report price changes within 30 minutes to a central aggregator under the Motor Fuel Price (Open Data) Regulations 2025. The government's Fuel Finder service then passes those prices, with forecourt locations and amenities, to authorised apps and websites. A comparison app, a navigation service or an advertiser running location-based campaigns all draw on the same feed.
From Sebastopol to the Open Data Directive
The modern meaning of the term was fixed on December 7-8, 2007, when 30 open government advocates met in Sebastopol, California. They agreed eight principles: public data should be complete, primary, timely, accessible, machine processable, non-discriminatory, non-proprietary and licence-free.
Governments followed quickly. Federal chief information officer Vivek Kundra launched data.gov in May 2009 with 47 datasets. The UK's data.gov.uk went live on January 21, 2010 with more than 2,500, backed by Sir Tim Berners-Lee. Berners-Lee and Sir Nigel Shadbolt founded the Open Data Institute in London in 2012 with a GBP 10 million public funding pledge. President Barack Obama's executive order of May 9, 2013 made "open and machine readable" the default for federal information, and the International Open Data Charter, adopted by 17 governments at the Open Government Partnership summit in Mexico on October 29, 2015, put "open by default" first among its six principles.
Statute arrived later. The US OPEN Government Data Act, Title II of the Foundations for Evidence-Based Policymaking Act, was signed on January 14, 2019 and turned the Obama-era policy into law. In Europe, Directive (EU) 2019/1024, the Open Data Directive, was adopted on June 20, 2019 and replaced the public sector information rules of 2003. It made reuse free of charge by default, extended scope to public undertakings and publicly funded research data, required dynamic data to be released "immediately after collection" via APIs, and created the category of high-value datasets. Member states had until July 17, 2021 to transpose it.
The implementing list followed in Implementing Regulation (EU) 2023/138, adopted on December 21, 2022 and applicable from June 9, 2024. It names six categories - geospatial, earth observation and environment, meteorological, statistics, companies and company ownership, and mobility - and requires them to be free, available through APIs and bulk download, and licensed under CC BY 4.0, CC0 or something less restrictive.
Why marketers keep running into it
Much of the location layer in advertising rests on open geodata. OpenStreetMap, started by Steve Coast in 2004 out of frustration with Ordnance Survey licensing, switched to the ODbL on September 12, 2012. The Overture Maps Foundation, formed by Amazon Web Services, Meta, Microsoft and TomTom in December 2022, reached general availability on July 24, 2024 with almost 54 million places of interest. Geofences, store locators, trade-area analysis and the business listings that feed local search all depend on accurate points of interest, and open sources compete directly with proprietary map providers for that role. Weather-triggered campaigns, which change creative or bids when temperatures cross a threshold, draw on meteorological data that the EU now classes as high-value.
Platforms publish open or public datasets of their own. Chrome's User Experience Report, explained in PPC Land's CrUX entry, covered 18,294,881 origins in its August 2026 release, and four experimental ad metrics put publishers' ad loads on the public record in September 2026. Schema.org and Google released usage statistics for 958 types and 4,587 properties as CSV and JSON files on a public GitHub repository. Google opened an alpha Trends API on July 24, 2025. Not all carry formal open licences.
The largest commercial use is now AI training. Common Crawl has published crawls of between 1 and 4 billion pages every few weeks since 2013, free to download. Wikipedia, licensed for reuse under Creative Commons, counts Amazon, Google, Meta, Microsoft and Mistral AI among paying Wikimedia Enterprise customers who want reliable, high-volume access to material they could legally take for nothing.
Quality, cost and privacy disputes
Open does not mean correct. An RAC Foundation check of Fuel Finder on March 30, 2026 found 368 open stations with coordinates more than 100 miles from their published postcode, 116 placed outside the UK landmass, and nine reporting a price below 2p per litre, according to the foundation.
Sustainability is a second fault line. Publishers bear the cost while reusers capture most of the value. Wikimedia said bots accounted for 65% of its most resource-heavy traffic in 2025 and linked a partial outage to agents it believes OpenAI operates. Political priorities also shift: more than 2,000 datasets were removed from data.gov after the change of US administration in January 2025.
Re-identification limits what can be released. Australia's health department published a 10% sample of Medicare and pharmaceutical benefits records covering about 2.5 million people on August 1, 2016. Researchers at the University of Melbourne showed that encrypted provider numbers could be decrypted, and the data was withdrawn on September 8, 2016, according to the Office of the Australian Information Commissioner. Anonymisation that looks adequate at publication can fail later.
Commercial reuse by the largest companies is contested too. Publishers argue that "free to read" was never "free to train on": the News/Media Alliance sent Common Crawl a demand letter dated April 29, 2026, noting that over 60% of its 2024 donations came from AI-affiliated sources. Common Crawl's material is publicly accessible web content rather than openly licensed data, a distinction the dispute turns on.
Not the same as
Open source applies openness to software code, governed by the Open Source Definition that PPC Land's explainer on open source traces to 1998. Prebid is Apache-licensed software; it publishes no datasets.
Open standards are freely implementable specifications. OpenRTB defines how bid requests are structured, but the bidstream data flowing through it is private, commercially traded and often personal.
Open banking is a mandated, consent-based channel through which banks share an individual customer's account data with authorised third parties. It is closer to data portability than to open data, because access is restricted to named parties.
Mandated data sharing and clean rooms also differ. Rivals receiving Google search logs under competition law get access on set terms, not a public release, and clean rooms exist to let parties match data without anyone seeing the other side's records.
Recent developments
On November 19, 2025, the European Commission's Digital Omnibus proposed folding the Open Data Directive and the Data Governance Act into the Data Act, and allowing public bodies to set higher fees and special conditions for "very large enterprises". The Dutch government supported a single data regulation while warning about weaker privacy protections. The COMMUNIA Association, an advocacy group for the public domain, warned on February 27, 2026 that the changes "would likely prompt public sector bodies to abandon open licences". As of October 2026 the proposal sits with the Parliament's ITRE and LIBE committees.
Regulators are also using openness as a competition remedy. On July 16, 2026 the Commission ordered Google to share anonymised search data with eligible rivals from January 2027, after rejecting a method that would have stripped 90% to 100% of unique queries. In the US ad tech case, Judge Leonie Brinkema's September 2, 2026 order rejected open-sourcing the auction logic of Google's publisher ad server, preferring conduct rules and per-ad data files. Meanwhile the European Data Protection Board adopted guidelines on July 7, 2026 on scraping personal data from public websites for AI training, a reminder that public visibility and open licensing are separate questions.
Timeline
- 2004 - Steve Coast starts OpenStreetMap in the UK.
- December 7-8, 2007 - 30 advocates meeting in Sebastopol, California agree the eight principles of open government data.
- May 2009 - The US launches data.gov with 47 datasets.
- June 29, 2009 - Open Data Commons releases the Open Database License 1.0.
- January 21, 2010 - The UK launches data.gov.uk with more than 2,500 datasets.
- 2012 - Sir Tim Berners-Lee and Sir Nigel Shadbolt found the Open Data Institute.
- September 12, 2012 - OpenStreetMap switches to the ODbL.
- May 9, 2013 - A US executive order makes open, machine-readable data the federal default.
- October 31, 2014 - The UK's Open Government Licence 3.0 is published.
- October 29, 2015 - 17 governments adopt the International Open Data Charter.
- August 1 to September 8, 2016 - Australia publishes, then withdraws, a Medicare and pharmaceutical benefits sample after re-identification research.
- January 14, 2019 - The US OPEN Government Data Act is signed.
- June 20, 2019 - The EU adopts the Open Data Directive (EU) 2019/1024.
- July 17, 2021 - Transposition deadline for the Open Data Directive.
- December 2022 - AWS, Meta, Microsoft and TomTom form the Overture Maps Foundation.
- December 21, 2022 - The Commission adopts the high-value datasets regulation (EU) 2023/138.
- June 9, 2024 - High-value dataset obligations apply across the EU.
- July 24, 2024 - Overture Maps reaches general availability with almost 54 million places.
- January 2025 - More than 2,000 datasets are removed from data.gov.
- November 19, 2025 - The Digital Omnibus proposes merging the Open Data Directive into the Data Act.
- February 2, 2026 - The UK's mandatory Fuel Finder open data scheme starts.
- July 16, 2026 - The Commission orders Google to share anonymised search data with rivals.
Related PPC Land coverage
- Explaining CrUX - Chrome's public dataset of real-user performance and its BigQuery tables.
- Chrome puts publishers' ad loads on public record with 4 new metrics - Experimental ad metrics added to the public CrUX dataset.
- Google and Schema.org finally show how the web uses structured data - Monthly usage statistics for Schema.org types and properties published on GitHub.
- Google opens alpha testing for new Trends API - Programmatic access to a rolling 1,800-day window of search interest data.
- Common Crawl supplies paywalled content to AI companies despite publisher objections - How the free web archive feeds AI training.
- News publishers target Common Crawl, the AI training data backdoor - The News/Media Alliance's April 2026 demand letter.
- Wikipedia turns 25 with eyes on AI survival and tech partnerships - Wikimedia Enterprise partners and the cost of AI reuse.
- Wikimedia says OpenAI agents may have contributed to partial outage - Bot traffic and infrastructure strain on an open knowledge project.
- Explaining open source - The Open Source Definition and its role in ad tech.
- Explaining Prebid - The Apache-licensed header bidding software.
- Explaining OpenRTB - The IAB Tech Lab protocol for real-time bidding.
- Explaining clean room - Controlled data matching as the opposite model to open release.
- Europe's Data Act reshapes connected device rules for marketers - The EU law set to absorb the Open Data Directive.
- Netherlands raises serious concerns about EU Digital Omnibus privacy changes - The Dutch position on merging EU data laws.
- Brussels forces Google to hand rivals its search data - The July 2026 DMA decision on search data sharing.
- Google loses fight to strip 90%+ of unique queries from EU rivals' data - The anonymisation method the Commission imposed.
- Google faces six-year worldwide ad tech decree instead of AdX sale - The ruling that rejected an open-source auction remedy.
- EDPB blocks AI firms from using consent as an excuse to scrape - Guidelines on scraping public web data for AI training.
Summary
Who. Governments and public bodies publish most open data, joined by volunteer projects such as OpenStreetMap and Wikipedia, nonprofits such as Common Crawl and corporate consortia such as Overture Maps. Developers, researchers, AI companies, ad tech vendors and advertisers reuse it. The Open Knowledge Foundation defines the term, and the EU, UK and US legislate for it.
What. Data that anyone can access, use, modify and share for any purpose, released under an open licence such as CC0, CC BY 4.0, the ODbL or the OGL, in machine-readable formats, through bulk downloads or APIs, at no more than marginal cost.
When. The eight Sebastopol principles date from December 2007. National portals followed in 2009 and 2010, the US OPEN Government Data Act in January 2019, the EU Open Data Directive in June 2019 and EU high-value dataset obligations in June 2024. A November 2025 proposal would merge the EU rules into the Data Act.
Where. Worldwide, through portals such as data.gov and data.europa.eu, with the most detailed legal framework in the European Union and the United Kingdom.
Why. Public and pooled data gains value when anyone can build on it. For marketing, open geodata, weather, mobility and price feeds underpin location targeting and local search, while openly available web content has become the raw material of AI systems - which is why licensing, quality and who pays for publication are now contested.
Discussion