Your Data Lake Is Probably Becoming a Data Swamp

July 21, 2026  |  Mark Hillary

Play Audio Version

How much data does your company store? It may be a lot more than you think and this requirement to store and retrieve data is becoming more and more important as data regulations become stronger and consumers demand more from the companies they interact with.

The figures for YouTube are almost certainly more than most corporate environments, but they give a clear picture of the problem. Over 20 million new videos are uploaded to YouTube every single day. Based on average video lengths the approximate amount of new content that is created each day is 82 years of video.

Two weeks from now, there will be over 1,000 years of new video on YouTube that was not there a fortnight ago. Just imagine the storage and data retrieval complexity involved in managing this.

Your company is unlikely to rival YouTube, but think about the data your business is capturing – especially in the interaction with your customers. You almost certainly have video and audio recordings of customer calls, chat transcripts, CRM history, QA scores, agent performance data, all the interaction metadata, call routing logs, customer journey data, and AI-generated summaries – just to make sense of it all.

Recording large amounts of video and audio is where the biggest challenge for storage is found. If your business only stored text-based data then several years of transaction data run to a few terabytes, but video multiplies this many times.

Data storage needs serious planning because there are different types of data and different uses. Some data must be recorded to meet industry regulations, but it is unlikely to be needed often. This can be stored away and made available when requested – with a short delay. Some data needs to always be immediately available.

This is why a data lake can be an important business asset for most companies. It’s a very flexible and scalable approach to strategically manage how you store different types of data and takes into account the need for data governance, analysis, enrichment, or the need to be used as AI training data.

A data lake is different to a traditional data warehouse. A data warehouse needs data to be cleaned up and structured. It is an ordered store of data that conforms to a certain standard. A data lake can contain anything. Structured data, unstructured data, logs, recordings, video, images, PDFs – you can drop anything in there.

This doesn’t mean that data lakes are just chaotic (but they can be) – they are usually divided into zones that helps to give some organization.

  • The raw – or landing – zone is where data arrives exactly as it is with no formatting or cleaning.
  • The clean – or curated – zone is where tagged, standardized, or data with duplicates removed resides.
  • The analytics zone is where you will find data that has been prepared for machine learning or reporting.
  • The archive – also known as cold storage – is where data that is not required often goes. It is stored and can be retrieved, but data in this zone is never required immediately.

The sheer scale of data lakes can lead to problems. It can be hard to query a data lake effectively when there is such a huge volume of data in various formats. The beauty of a database formatted into tables is that you can build complex queries that return results extremely quickly.

Another major issue is when a data lake becomes a data swamp. You keep adding more and more data, paying for all that storage, and losing sight of what is not needed and can be removed. This can become expensive, especially when you don’t know what can be removed because you have lost track of what is in there.

The best way to manage a data lake is to start with your business problem – what is it you need from the data? You can then define what needs to be stored, what you are allowed to store, and how best to organize it to allow business intelligence and ongoing maintenance.

It is easy to sign a contract with a cloud provider such as AWS or Azure that gives you access to unlimited storage, but building a data lake that facilitates solid business intelligence requires planning. It also needs a strategy for ongoing maintenance so you are not wasting budget paying to store unneeded data forever.

Even a medium-sized company will now be storing terabytes of data for their business processes. Larger companies will be storing petabytes worth of data. Once you get into companies with extensive AI and IoT requirements then the storage requirements may be up in the exabyte range. To put this into context, 5 petabytes could easily store every printed word in every library in the world.

The most important lesson is that data storage is no longer just an IT infrastructure question. You can no longer just leave storage decisions for the CIO. This is now a strategic design question that will affect how your business can function – effectively or not.

What should be stored, where it should stay, how quickly it may need to be retrieved, how long it should be retained, and how it can be used for analytics or AI all need to be planned before the technology decisions are made.

A poorly designed data lake can quickly become a very expensive data swamp, but a well-designed one can become one of the most valuable assets in your business.

This is why companies need experienced advice before making major storage decisions. The right advisor will help define the business problem first, then design the data architecture around governance, compliance, cost, retrieval speed, and future value.

Storage is easy to buy. Intelligent data design is much harder – and far more important for any business that relies on business intelligence to function.

For more information on data lake design and management with IBA Group please click here.

    Access full story Leave your corporate email to get a file.
    Yes

      Subscribe A bank transforms the way they work and reach
      Yes