Companies use public data to assess business partners, assess risk, plan transport, select locations and forecast demand. Registers, statistics, maps and weather data are no longer merely material for occasional analysis. Increasingly, they are integrated into products and operational processes.
This has changed the nature of the relationship. An error in a public register can influence a credit decision. A delay in updating transport data reduces the quality of an application. A change to the API format can halt an automated process. The source remains public, but the risk becomes commercial.
Open data is just the beginning
From 9 June 2024, high-value datasets in the European Union must be made available free of charge, in their current version, in machine-readable format, via an API and for bulk download. The regulations cover geospatial, environmental, meteorological, statistical, transport and business data.
The regulation has lowered the barrier to access. However, it has not solved the fundamental problem for businesses: the difference between a published dataset and a reliable data product.
A data product has an owner, an update schedule, documentation, a change history and stable identifiers. Its users know when the response structure will change, how long previous versions will remain available and where to report an error. The API itself does not provide any of these features.
When these features are missing, the cost does not disappear. It is passed on to businesses. Each company individually maps names, removes duplicates, compares records, archives successive versions and protects its system against changes to the source. An ‘integration tax’ arises, which is not reflected in the price of the data.
The scale of usage reflects real demand
The GUS Local Data Bank is Poland’s largest database of information on the economy, society and the environment. Its API provides data via REST in JSON and XML formats, has public documentation, defined limits and a channel for handling technical issues.
Since December 2018, the API has handled over 280.9 million queries. It is used by nearly 9,400 registered users.
This is not a measure of the portal’s popularity. It is a sign that the public database has become an integral part of systems developed outside the public administration. The greater the usage, the greater the importance of the service’s stability, quality and predictability.
This is even more evident in the UK. In the financial year ending 31 March 2026, data from Companies House was accessed 14.6 billion times. The register covered 5.48 million entities. The UK government estimates its annual value to users at over £1–3 billion.
At the same time, Companies House illustrates why mere availability is not enough. Registry data is used for risk assessment, granting finance and verifying business relationships. Its quality therefore affects not only user convenience, but also the flow of capital and the ability to detect fraud. For this reason, the UK register is expanding identity verification, data checks and sanctions against entities providing unreliable information.
Most value is created before the API is made available
Transport for London demonstrates that public authorities can not only publish data but also reduce the cost of using it. TfL has standardised information from various modes of transport into a single model. Developers do not need to create separate logic for the Underground, buses and other services. They receive data in a consistent structure, independent of the source systems.
Over 17,000 registered developers use the solution. TfL’s data powers more than 600 apps used by over 40 per cent of Londoners. The annual benefits and savings have been estimated at up to £130 million.
This value did not arise simply from the number of files published. It arose because TfL carried out some of the work that would otherwise have had to be repeated numerous times by hundreds of companies.
A public data source requires a designated owner within the organisation
Before integrating public data into a key business process, a company should define its role just as precisely as it would for an external technology provider. Who is responsible for quality? How are structural changes detected? Can a previous state be restored? What will the system do if the source stops responding?
This is not solely a matter for the IT team. The quality of the source affects revenue, risk, compliance and customer service. The decision to use public data is therefore a decision about a business dependency.
Public data delivers the greatest value when the state takes responsibility for standardisation and reliability, whilst companies compete on the quality of the services built upon it. When the burden of preparing the data source falls on the market, businesses end up paying repeatedly to solve the same problem.
The economic significance of public data is not determined by the number of datasets made available. It is determined by the cost of integrating them into an existing product.

