Every data platform roadmap eventually runs into the same wall: a warehouse or lake that locks data inside one vendor’s engine, one file layout, one query language. An open table format removes that wall by adding a shared metadata layer that any engine, Spark, Snowflake, Trino, can read and write against the same underlying files.

For CEOs and CTOs, the real question is who controls the roadmap for the data behind your AI, analytics, and reporting, and how much it costs to change course later. Discover the answers in our comprehensive guide.

Key takeaways

  • These open formats add a shared metadata and transaction layer on top of Parquet-style files, so multiple engines can read and write the same data without duplicating it.
  • Apache Iceberg, Delta Lake, and Apache Hudi all trace back to a real production problem at Netflix, Databricks, and Uber, each of which hit the limits of Hive tables at scale.
  • Migrating off proprietary storage reduces vendor lock-in and lets engineering teams run different compute engines against one shared copy of the data.
  • A staged migration, assess, pilot, backfill, cutover, optimize, keeps the business running while the platform changes underneath it.
  • Format choice increasingly overlaps with catalog choice. Databricks’ acquisition of Tabular and Snowflake’s decision to open source its Polaris Catalog both point toward the market consolidating around Iceberg compatibility.
  • The hardest part of a migration is rarely the file conversion itself. It is governance, query rewrites, and keeping downstream BI tools working through cutover.

What is an open table format?

An open table format is a specification layered on top of columnar files such as Parquet or ORC. It adds atomic commits, schema evolution, partition evolution, and time travel across historical snapshots, using a metadata layer that tracks which files belong to which version of the table.

Three formats dominate the space, and each grew out of a real production challenge. Netflix developed Apache Iceberg to address scalability, correctness, and table-evolution limitations in Hive-based data lakes. Uber built Apache Hudi in 2016 to support low-latency incremental ingestion, updates, and deletes across its large data lake. Databricks developed Delta Lake to bring ACID transactions, scalable metadata, and reliable batch and streaming processing to data lakes built around Spark. Delta Lake was open sourced in 2019 and moved under Linux Foundation governance later that year.

The result behaves like a database table, atomic writes, concurrent readers, rollback, while living on ordinary cloud object storage. That combination is what makes open table formats the foundation of the modern lakehouse.

Why leadership teams are migrating to open table formats now

For most of the last decade, the practical choice was to pick a cloud data warehouse and accept that your data lived inside its proprietary format. That trade-off is no longer necessary, and it is getting more expensive to keep making.

Three forces are pushing migration up the executive agenda:

  • Negotiating position. When your data sits in a proprietary format, switching compute engines means re-exporting everything. Adopting an open standard lets you keep one copy of data and point different engines at it, which changes your position at contract renewal.
  • Consolidation around compatibility. Databricks acquired Tabular, founded by Apache Iceberg’s original creators, in 2024 in a deal reported at more than $1B. Databricks framed the acquisition around improving compatibility between Iceberg and Delta Lake. Around the same time, Snowflake introduced Polaris Catalog, an open catalog for Apache Iceberg, and open sourced it in July 2024. Together, these moves point to interoperability and Iceberg compatibility becoming increasingly important across the major data platforms.
  • AI and analytics workloads outgrowing single-engine setups. Training and inference pipelines, BI dashboards, and streaming jobs increasingly need to read the same tables without a copy step in between.

N-iX data engineers see this pattern across clients. The trigger for migration is rarely a single incident. It is the accumulation of small workarounds that finally costs more than the move itself.

Choosing the right open table format for migration

Format interoperability tools like Apache XTable and Delta Lake’s UniForm now translate metadata between Iceberg, Delta Lake, and Hudi, so this decision is reversible in ways it wasn’t a few years ago. Still, most teams pick one format as their system of record, and the decision usually comes down to workload shape and ecosystem fit.

The table below summarizes the factors that tend to decide the question for enterprise teams.

Decision factor

Apache Iceberg

Delta Lake

Apache Hudi

Best fit

Batch-heavy analytics, multi-engine access

Spark-centric pipelines, Databricks ecosystem

High-frequency upserts and deletes (CDC, streaming)

Governance

Apache Software Foundation

Linux Foundation

Apache Software Foundation

Catalog ecosystem

Broadest vendor support (Snowflake, AWS, Databricks, Google)

Strongest inside Databricks, expanding via UniForm

Growing, historically tied to Spark tooling

Origin

Built at Netflix

Built at Databricks

Built at Uber

N-iX experts highlight: Governance matters as much as features. Apache Iceberg and Apache Hudi are governed by the Apache Software Foundation. Delta Lake is governed by the Linux Foundation, the same nonprofit that stewards Kubernetes and Node.js. None of the three formats answers to a single vendor’s product roadmap, which is the property most enterprises are actually migrating to secure.

How to migrate to an open table format in 6 stages

A migration that treats the format switch as a single weekend cutover is the fastest way to break downstream reporting. The staged approach below keeps the current platform live while the new one proves itself in production.

1. Assess and scope

Inventory which tables genuinely need format-level features, ACID writes, concurrent readers, time travel, and which are stable enough to leave alone. Not every table justifies a migration.

2. Pick a target format and catalog

Match the choice to workload shape using the comparison above, and settle on a catalog, Unity Catalog, Polaris, AWS Glue, Hive Metastore, before writing any pipeline code.

3. Run a bounded pilot

Convert one or two non-critical tables end to end, including the dashboards and jobs that read them, and measure query performance and cost against the current setup.

4. Build the backfill and dual-write bridge

Stand up a pipeline that writes new data into both old and new tables while historical data is backfilled in batches, so nothing downstream breaks mid-migration.

5. Cut over with validation gates

Switch reads to the new tables one at a time, checking row counts, query results, and job success rates at each gate before moving to the next table.

6. Optimize and decommission

Tune compaction, partitioning, and file sizes for the new format, retire the dual-write path, and archive the legacy tables once every consumer has moved.

N-iX engineers who run these migrations treat the pilot stage as the one that gets skipped under deadline pressure, and it is the stage that prevents the worst surprises later.

Common obstacles in open table format migration

Most migration setbacks are predictable, and most of them show up in the same three places.

  • Small file storms during backfill. Naive backfills that copy history table by table generate huge numbers of tiny files, which slows every query that follows. N-iX engineers stage backfills in batches sized to the new format’s target file size, then run compaction before opening the table to production reads.
  • Catalog fragmentation. Teams that migrate the file format without settling on one catalog end up with tables that different engines discover differently. We settle the catalog decision before migration starts, before the first broken pipeline forces the issue.
  • Skills gaps slowing adoption. Engineers fluent in a warehouse’s proprietary SQL dialect often need real time to get comfortable with table format internals such as partition evolution, snapshot expiration, and metadata compaction. Our engineers pair migration work with hands-on enablement so your platform team can run the new tables independently once we hand the system back.

How N-iX helps you migrate to an open table format

With more than 2,400 technology professionals across our delivery centers, we can staff a migration with the platform, pipeline, and governance specialists it actually needs, so no single generalist team gets stretched across every stage.

Security and compliance questions come up early in any data migration, especially when the tables involved hold regulated or customer data. N-iX holds SOC 2 Type 2 certification and has maintained ISO 27001 certification since 2017. The controls around how we handle client data are independently audited each year.

We scope every open table format migration the way we describe above. Assess first, prove the pattern on a pilot, then expand once the numbers hold up. If your team is weighing this move, reach out to us and size the effort before you commit to it.

FAQ

Is this the same as migrating to a new data warehouse?

No. A table format change affects how data is stored and versioned underneath your existing engines. Most migrations keep the same warehouse or lakehouse platform in place and swap the storage layer beneath it, so the process can run alongside normal operations without a full platform cutover.

How long does a typical migration take?

It depends on the number of tables, how many downstream jobs and dashboards read them, and how much historical data needs backfilling. A pilot on one or two tables can run in weeks. A full enterprise rollout across hundreds of tables is measured in months, staged the way we describe above.

Do we have to choose one format permanently?

Not anymore. Tools like Apache XTable and Delta Lake’s UniForm translate metadata between Iceberg, Delta Lake, and Hudi, so most organizations can change course later without a second full migration. Most teams still pick one format as their primary system of record to avoid running duplicate pipelines.

Are these formats only useful for AI and analytics workloads?

No. While AI training and inference pipelines are a common driver, open table formats are equally valuable for standard BI reporting and financial consolidation. The same guarantees apply to any workload where multiple teams need concurrent, consistent access to the same data.

What is the biggest risk during migration?

The biggest risk is downstream breakage. A dashboard, scheduled job, or partner feed that depends on the old table structure can fail silently during cutover. Staged validation gates, checking row counts and query results before moving to the next table, catch this before it reaches production users.

Does a migration require replacing our existing data team?

No. The staged process above is designed to run with your existing platform and data engineering staff. Outside specialists fill specific gaps, catalog decisions, performance tuning, governance setup, and hand the platform back to the team that will operate it.

Have a question?

Speak to an expert
N-iX Staff
Valentyn Kropov
Chief Technology Officer

Required fields*

Table of contents