Data engineer reviewing open table format architecture and data pipeline dashboards at a modern workstation.

Snowflake, Databricks, and AWS All Just Bet on the Same Open Table Format

Fri, Aug 14, 2026

Skip the “what is a lakehouse?” explanation. You have probably read it already, possibly in Refonte Learning’s existing data-engineering coverage, where Delta Lake, Apache Iceberg, and Hudi already appear alongside ACID transactions and schema evolution.

What matters in Apache Iceberg 2026 is not that another table format exists. What matters is that Snowflake, Databricks, AWS, and Confluent are increasingly designing their catalog, storage, governance, and streaming interfaces around Iceberg-compatible standards, turning what used to look like a three-way format contest into a more complicated consolidation around Iceberg at the interoperability boundary.

There is an important date correction before we go further. Not every component of this consolidation originated in 2026: Amazon S3 Tables launched in December 2024 and added Iceberg REST Catalog-compatible APIs on March 13, 2025, while Confluent made Tableflow for Iceberg generally available in March 2025. What changed by 2026 is that these earlier bets lined up with much more explicit 2026 moves from Snowflake and Databricks, Iceberg v3 reached an adopted specification, and vendors began treating the Iceberg catalog interface as infrastructure rather than an optional compatibility feature.

That distinction matters. A vendor blog saying “we support Iceberg” is not the same signal as a storage service exposing the Iceberg REST Catalog API, a warehouse vendor restructuring its managed Iceberg catalog around that interface, or a streaming platform materializing Kafka topics into Iceberg tables.

This is the open table format consolidation story: the dated moves, the technical boundaries that actually matter, what the Iceberg v3 spec changes, where the vendor claims deserve skepticism, and what all of it means for data engineer lakehouse skills in 2026.

The Open Table Format Landscape Nobody Was Consolidating... Until 2026

For years, Apache Iceberg, Delta Lake, and Apache Hudi coexisted as credible open table formats. The meaningful question for an architecture team was often which implementation fit its compute engine, mutation pattern, governance stack, and operating model best.

That landscape has not disappeared. What has changed is where vendors are investing interoperability effort: Snowflake has embedded Apache Polaris into Horizon Catalog for Iceberg REST interoperability, Databricks made managed Iceberg and Iceberg v3 generally available in Unity Catalog in May 2026, and AWS continues to describe S3 Tables as fully managed Apache Iceberg tables rather than as a format-neutral table-storage service.

Confluent provides a fourth signal from a different part of the pipeline. Tableflow materializes Kafka topics as Iceberg or Delta tables, but its Iceberg path comes with a built-in Iceberg REST catalog and external catalog integrations, making open table metadata part of the streaming-to-analytics handoff.

Vendor move

What it signals

Snowflake directs new Open Catalog customers toward Horizon Catalog

Snowflake is actively restructuring its Iceberg catalog architecture, not merely adding an Iceberg file reader

AWS builds S3 Tables specifically around Apache Iceberg

The object-storage layer itself now understands managed Iceberg tables and exposes an Iceberg-compatible catalog interface

Databricks makes Managed Iceberg GA through Unity Catalog

Databricks is no longer positioning open interoperability solely through Delta Lake plus UniForm

Confluent materializes Kafka topics as Iceberg tables

The streaming layer can publish an analytics-ready open-table representation without a separately operated Kafka-to-lake ETL pipeline

Iceberg v3 becomes an adopted specification

Row lineage, deletion vectors, new types, defaults, transforms, and encryption-key metadata move into the stable format contract

The phrase “Iceberg vs Delta Lake 2026” therefore needs a different interpretation from the same comparison three years ago. The question is becoming less “which single format will own my entire stack?” and more “which format is native to my platform, which interfaces are open at the catalog boundary, and how reversible is the choice?”

That is a better architecture question because the physical table representation is only one layer. The catalog has become the control point through which engines discover metadata, obtain credentials, enforce policy, coordinate writes, and access the current snapshot.

Why the Site's Existing "Big Data Trends" Coverage Doesn't Cover This

Refonte Learning’s guide to the essential data engineering tools already covered on this site, including Kafka and Airflow, introduces Iceberg, Delta Lake, and Hudi as part of the broader move toward reliable lakehouse tables. That foundation is useful, but it is deliberately different from the catalog-level consolidation happening now.

The same distinction applies to the site’s 2026 trends coverage. Iceberg and Delta appear there as technologies within a much larger data-engineering landscape; Unity Catalog, the Iceberg REST interface, Snowflake’s Horizon transition, and the production implications of Iceberg v3 are not the organizing thesis.

That gap is precisely where this article sits:

  • It does not rebuild an ACID-transactions primer.

  • It treats table format and catalog as separate architectural decisions.

  • It distinguishes launches that really happened in 2026 from older products whose importance became clearer in 2026.

  • It treats claims such as “10x throughput,” “zero ETL,” and “no lock-in” as vendor claims that need architectural interpretation, not conclusions.

That last point is important. Open specifications reduce specific forms of lock-in; they do not eliminate cloud IAM, proprietary governance policies, optimizer behavior, operational tooling, egress costs, or engine-specific features.

Snowflake Quietly Restructured Its Iceberg Catalog Strategy

Snowflake’s current Open Catalog documentation contains one of the clearest signals in the market. It says customers that have never created a Snowflake Open Catalog account can no longer sign up for their first one and directs new customers to Snowflake Horizon Catalog for Iceberg tables and multi-engine interoperability.

Existing Open Catalog customers can continue using the service and, when eligible, create additional Open Catalog accounts. So this is not a shutdown or forced immediate migration.

Snowflake also now recommends Horizon Catalog rather than Open Catalog synchronization when third-party engines need to access Snowflake-managed Iceberg tables. Its documentation explicitly says synchronizing those tables with Open Catalog is “no longer recommended” because Horizon can expose them directly.

The technical reason is more interesting than the product naming. Snowflake integrated Apache Polaris into Horizon Catalog and made external-engine querying of Snowflake-managed Iceberg tables generally available on February 6, 2026 through the Iceberg REST protocol.

Snowflake followed that with Iceberg v3 support in March 2026 and described Horizon Catalog as exposing tables through a standardized Iceberg REST interface powered by Polaris.

What Snowflake says now

Architecture implication

First-time Open Catalog signups are closed

New designs should not assume Open Catalog is Snowflake’s strategic onboarding path

Existing customers may continue

There is no evidence of an immediate forced deprecation

Horizon is recommended for multi-engine access

Iceberg interoperability is moving into Snowflake’s broader governance/catalog control plane

Polaris powers Horizon’s interoperability layer

Open REST interfaces now sit closer to Snowflake’s primary catalog architecture

External Iceberg querying reached GA in February 2026

The interoperability path is production-positioned, not merely conceptual

There is one documentation issue a senior engineer should not ignore. The Open Catalog page still says billing “will begin in the first half of 2026”, even though the current date is August 14, 2026.

That makes the page internally stale on timing. Treat the statement as evidence of Snowflake’s intended billing transition, not as proof that billing actually began on a particular date; I could not verify a dated Snowflake release note that establishes the actual commencement day.

That is exactly why procurement and architecture decisions should use current service-consumption documentation and account-specific commercial terms rather than copying a billing sentence from a product overview.

What Horizon Catalog Changes for Snowflake Customers

The practical change is not simply “Open Catalog was renamed.” Horizon Catalog is a wider governance and metadata layer, while Open Catalog was Snowflake’s managed Apache Polaris service specifically oriented toward Iceberg cataloging. Snowflake’s Horizon documentation describes a catalog spanning data inside and outside Snowflake with governance, lineage, quality, and cross-engine interoperability.

For a new 2026 architecture, that means the evaluation boundary changes.

  • New Snowflake customer: evaluate Horizon’s Iceberg REST path first because Snowflake directs first-time users there.

  • Existing Open Catalog customer: do not assume forced migration, but map dependencies and ask what functions should remain in Open Catalog versus move into Horizon.

  • External-engine user: test your actual Spark, Trino, or other client against Horizon’s supported authentication, credential vending, write semantics, and policy behavior rather than stopping at “REST compatible.”

  • Cost-sensitive team: verify current billing directly; even Horizon’s REST API documentation says its own billing is scheduled for the second half of 2026 and remains subject to change.

The bigger lesson is that a catalog service has its own lifecycle. An open table format does not protect you from operational changes in the managed catalog controlling access to it.

Databricks Bet on Interoperability, Not Lock-In

The Databricks part of this story deserves precise dating because one milestone is repeatedly mischaracterized as 2026 news.

Databricks open-sourced Unity Catalog under Apache 2.0 on June 13, 2024, hosted through LF AI & Data. That announcement also documented compatibility with the Hive Metastore API, implementation of the Iceberg REST Catalog API, and Iceberg access through UniForm.

That event predates 2026 by almost two years. It should be treated as background to Databricks’ interoperability strategy, not recycled as a new 2026 announcement.

The real 2026 move is stronger: on May 28, 2026, Databricks announced Managed Iceberg, Iceberg v3, and Foreign Iceberg as generally available in Unity Catalog. Managed Iceberg tables can be created, read, and written by external engines through Unity Catalog’s Iceberg REST Catalog APIs.

Current Databricks documentation says the Iceberg REST endpoint supports reads and writes for Managed Iceberg, reads for Foreign Iceberg, and Iceberg reads for appropriately configured managed or external Delta tables. Supported clients include Spark, Flink, and Trino.

That changes the Iceberg vs Delta Lake 2026 discussion considerably.

Question

Delta Lake on Databricks

Iceberg on Databricks in 2026

Is it native to Databricks?

Yes

Yes for Managed Iceberg

Can Unity Catalog govern it?

Yes

Yes

Can external Iceberg clients access it?

Yes, when Iceberg reads/UniForm are enabled

Yes through the Iceberg REST API

External write support

Depends on API/table mode

Managed Iceberg supports REST-based external writes

Primary format advantage

Deep Databricks-native feature integration

Broader direct alignment with the Iceberg client ecosystem

Does one format require copying Parquet data to expose the other?

UniForm can generate Iceberg metadata without rewriting the underlying Parquet files

Not applicable when the table is already Iceberg

UniForm remains important. Databricks documents that Iceberg reads for Delta tables generate Iceberg metadata alongside Delta metadata without rewriting the Parquet data files, so a single physical dataset can serve Delta-aware and Iceberg-aware clients.

That is not the same as saying Delta and Iceberg have become identical. Metadata semantics, feature support, writers, optimization behavior, and engine compatibility still differ.

What has changed is the cost of choosing between them. Databricks now has a stronger incentive to make Unity Catalog the governance layer regardless of whether the underlying managed table is Delta or Iceberg.

Databricks’ ecosystem argument also has real breadth, although its origin again dates to 2024: AWS, Microsoft Azure, Google Cloud, NVIDIA, Salesforce, DuckDB, LangChain, dbt Labs, Fivetran, Confluent, and other organizations supported the open-source Unity Catalog announcement.

The practitioner takeaway is not “Databricks abandoned Delta Lake.” It plainly has not.

The stronger conclusion is that Databricks no longer needs Delta-only access to preserve the value of Unity Catalog. The catalog can increasingly remain the strategic control plane even when an organization chooses Iceberg as its table representation.

AWS Built S3 Tables Directly on the Iceberg Standard

Amazon S3 Tables may be the cleanest evidence that Iceberg moved below the analytics-engine layer.

AWS describes S3 Tables as storage built on the Apache Iceberg open standard, with automated table maintenance, compaction, governance integrations, and an Iceberg REST Catalog API that works with Iceberg-compatible engines including Spark, Trino, Flink, Athena, Redshift, Snowflake, and third-party tools.

But again, get the chronology right. AWS launched S3 Tables on December 3, 2024; it added Iceberg REST Catalog-compatible table-management APIs on March 13, 2025.

So the claim “AWS launched S3 Tables on Iceberg in 2026” would be wrong.

The 2026 significance is that AWS now treats managed Iceberg as part of S3’s long-term identity. In a March 13, 2026 retrospective on S3’s twentieth anniversary, AWS highlighted S3 Tables specifically as “fully managed Apache Iceberg tables” alongside newer native S3 data capabilities.

AWS also expanded the surrounding operating model before 2026 with replication and Intelligent-Tiering, announced December 2, 2025. Those features automate cross-region/cross-account table replicas and tier table data based on access patterns.

S3 Tables capability

What I would evaluate in production

Native Iceberg REST Catalog compatibility

Whether every required engine supports the exact REST operations and auth model you use

Automated compaction

Whether AWS’s maintenance policy matches workload-specific file-size and clustering requirements

Spark/Trino/Flink interoperability

Read and write behavior, commit concurrency, credential handling, and feature-version support

Athena and Redshift integration

Query semantics and support for the specific Iceberg spec version your writers emit

Intelligent-Tiering

Access-pattern fit, minimum durations, operational savings, and workload predictability

Cross-region replication

Recovery objectives, replication lag, data-transfer cost, and metadata consistency

AWS currently advertises up to 10x higher transactions per second than Iceberg tables in general-purpose S3 buckets and up to 80% lower storage costs through storage-cost optimization. Those are AWS’s own performance and cost claims, not independent benchmark results.

Treat the wording “up to” seriously.

A 10x transaction-throughput ceiling tells you nothing about your median workload until you reproduce the comparison with the same writer count, commit frequency, partition layout, data-file size, catalog behavior, and maintenance policy. The same applies to the 80% storage figure: achievable tiering economics depend heavily on how much data becomes cold and how often queries touch it.

There is another compatibility detail engineers should watch. Current AWS documentation for EMR 7.12.0 notes that Athena SQL cannot read an Iceberg v3 table created by EMR Spark in that documented configuration because Athena reports v3 as unsupported.

That example is a useful antidote to “open standard means universal compatibility.” The spec can be adopted before every engine implements every format version, so compatibility matrices matter more than logos.

Confluent Is Turning Kafka Streams Straight Into Iceberg Tables

Confluent attacks the problem from the opposite direction: not storage-to-query, but stream-to-table.

Confluent Tableflow materializes Kafka topics as Apache Iceberg or Delta Lake tables. Its current documentation says Tableflow automatically handles schematization, type conversion, schema evolution, CDC materialization, and catalog publishing, reducing the need to operate a separate pipeline whose only job is continuously turning Kafka records into analytical tables.

Confluent calls this “zero-ETL.” Its July 10, 2026 explanation defines that phrase as eliminating custom ingestion/conversion/catalog jobs because Tableflow manages those stages for you.

As an engineer, I would translate “zero-ETL” more narrowly: zero separately managed ETL for the Kafka-topic-to-table materialization step.

You still have transformations somewhere if business semantics require cleaning, enrichment, deduplication, joins, privacy rules, aggregates, or dimensional models. Removing an infrastructure job does not remove data engineering.

Traditional Kafka-to-lake path

Tableflow-managed path

Kafka topic

Kafka topic

Sink/connect job

Tableflow

Object-storage landing

Managed by Tableflow

Parquet conversion

Managed by Tableflow

Iceberg metadata commits

Managed by Tableflow

Schema propagation

Integrated with Confluent’s schema machinery

Catalog publication

Built-in Iceberg REST catalog or supported external catalog integration

Downstream engine

Snowflake, Databricks, BigQuery, Trino, or other supported consumer

Confluent made Iceberg Tableflow generally available in Confluent Cloud on March 19, 2025, not 2026. Its original Tableflow announcement goes back to March 2024.

Where the 2026 timeline becomes interesting is Confluent Platform 8.3, announced July 29, 2026. The release explicitly describes governed, structured streaming data as usable across features “from Flink for real-time stream processing to Tableflow for unified lakehouse analytics.”

However, I would not state that Confluent Platform 8.3 “bundles Tableflow” into the self-managed platform. I could not verify that wording in Confluent’s official 8.3 announcement, and Confluent’s current Tableflow documentation explicitly describes Tableflow as a Confluent Cloud feature.

The defensible 2026 statement is narrower: CP 8.3 strengthens the governed streaming substrate that Confluent positions as feeding downstream Tableflow analytics, while Tableflow itself remains documented as a Confluent Cloud capability.

That distinction may sound pedantic. It is exactly the distinction that prevents an architecture team from buying against a feature diagram and discovering during implementation that the deployment model differs from what it assumed.

What Iceberg's V3 Spec Actually Adds: and Why the Timing Isn't a Coincidence

The Iceberg v3 spec is no longer a proposal. The Apache Iceberg specification currently states that versions 1, 2, and 3 are complete and adopted by the community, while version 4 remains under active development and has not been formally adopted.

That wording is more useful than tying v3 maturity to an arbitrary client-library point release. The Apache Iceberg GitHub release history and individual engine implementations move on separate schedules, so you should verify exact runtime support rather than infer it from the specification’s status.

Version 3 adds:

  • Nanosecond timestamp types, including timestamp-with-time-zone variants.

  • New types: unknown, variant, geometry, and geography.

  • Column default values.

  • Multi-argument transforms for partitioning and sorting.

  • Row lineage tracking.

  • Binary deletion vectors.

  • Table encryption keys tracked in metadata.

These are not all equally important for every workload. Geometry and geography matter to spatial analytics; variant matters for semi-structured data; nanosecond timestamps matter where microsecond precision is genuinely insufficient.

The broader maturity signal comes from the operational features.

Iceberg v3 addition

Why an engineering team cares

Row lineage

Preserves row identity and last-update sequence information across compatible updates

Binary deletion vectors

Makes positional deletes more efficient at execution than v2 position-delete files

Table encryption keys

Defines metadata structures for associating encryption-key information with tables/snapshots

Default column values

Makes schema evolution more expressive

Multi-argument transforms

Expands partitioning/sorting design options

Variant

Improves support for semi-structured values

Nanosecond timestamps

Supports higher-precision event and operational data

The timing does reinforce the vendor-consolidation story, although “not a coincidence” should not be read as evidence of a coordinated vendor plan.

A more defensible interpretation is that the forces reinforce each other: vendors have stronger incentives to build managed infrastructure around a specification as its production semantics mature, while broader vendor implementation gives the Iceberg community more pressure and feedback around security, row-level mutation, catalog behavior, and engine interoperability.

Row Lineage and Deletion Vectors: the Two Changes That Matter Most

Iceberg v3 row lineage introduces rowid, a unique long identifier assigned to a row, and lastupdated_sequence_number, which records the commit sequence that last updated the row. The specification defines how writers preserve those values when rows move between files and how readers inherit them from snapshot and file metadata.

One qualification matters: lastupdated_sequence_number is not itself a wall-clock “changed at” timestamp. It identifies the commit sequence; you use the surrounding snapshot/commit metadata to place that change in temporal context.

There is another limitation that marketing summaries tend to omit. The spec says row lineage cannot preserve lineage in the same way for updates performed through equality deletes because those writers do not read the original row and therefore cannot retain its original row ID.

That still makes row identity far more useful for debugging, change analysis, and governance-oriented workflows. But you should understand its mutation semantics before promising auditors “perfect row lineage.”

Deletion vectors deserve a similar correction.

Iceberg v2 did not rely only on copy-on-write. It already supported position and equality delete files. In v3, deletion vectors encode deleted positions for a particular data file as a bitmap; the specification explicitly describes them as more efficient at execution time than position-delete files and deprecates writers adding new v3 position-delete files.

That is an important production-hardening change.

Instead of scattering positional deletes across conventional delete-file records, an engine can use a compact bitmap representation associated with a specific data file. Readers still need correct implementation support, and writers must merge delete information so that a snapshot has at most one deletion vector per referenced data file.

The lesson: v3 improves row-level mutation mechanics, but do not enable format version 3 merely because the specification exists. Confirm that every engine reading or writing the table implements the features you intend to use.

What Data Engineers Actually Need to Do About This

No team needs an emergency lakehouse migration because four vendors have aligned around more Iceberg interfaces.

The rational response is evaluation. Start with the control points that can actually change your reliability, operating cost, or portability rather than migrating tables to participate in a trend.

For the larger context, Refonte Learning already covers the broader 2026 data engineering trends and tools landscape. The format-consolidation question is narrower: where can you delete custom infrastructure, where has the catalog contract changed, and which assumptions about interoperability need retesting?

Use this sequence:

  1. Inventory the catalog before the format. Record which service owns each table’s current metadata pointer, credentials, policies, and write coordination.

  2. Check Snowflake Open Catalog exposure. First-time signups are no longer available, while existing accounts remain supported; document which side of that boundary your organization occupies.

  3. Test REST compatibility end to end. “Implements Iceberg REST” matters more than a generic “supports Iceberg” badge, but only an actual read/write test proves your chosen auth and feature path.

  4. Benchmark S3 Tables against your current tables. Reproduce your own commit rates, query patterns, compaction needs, and storage-temperature profile before accepting AWS’s maximum savings figures.

  5. Identify Kafka-to-lake plumbing that exists only for materialization. That is where Tableflow can potentially remove operational work; business transformations will still exist.

  6. Build a format-version compatibility matrix. A v3 table is only as interoperable as the least-capable engine that must consume it, as the documented Athena/EMR example demonstrates.

The rule I use during migrations is simple: do not migrate a stable table because a destination is fashionable; migrate because the new architecture eliminates a measurable constraint.

That constraint could be proprietary metadata access, an ETL job you no longer need, a catalog that blocks another engine, expensive compaction operations, or governance duplicated across platforms.

If you cannot name the constraint, you probably do not yet have a migration case.

Skills Priority Order for Data Engineers Navigating the Format Wars

The most useful data engineer lakehouse skills in 2026 are not memorizing every table property. They are knowing which layer you are changing and what interoperability claim to test.

Priority

Skill

Must

Distinguish a table format such as Iceberg or Delta Lake from a catalog such as Horizon Catalog, Unity Catalog, or Apache Polaris

Must

Evaluate whether Iceberg support exposes a real REST Catalog-compatible interface and what read/write operations it supports

Must

Build a cross-engine compatibility matrix before enabling a new table-format version

Should

Operate or test one Iceberg-native path outside your primary platform, such as PyIceberg, Spark, S3 Tables, or Tableflow

Should

Separate vendor benchmark claims from your own workload measurements

Good

Understand v3 row lineage, deletion-vector, type, and encryption metadata semantics

Good

Design a catalog migration or federation path without unnecessarily rewriting table data

This is also where the essential data engineering platforms for 2025 careers become useful background: engines, orchestration, storage, and catalogs have to be understood as a system rather than as independent résumé keywords.

Table format versus catalog ranks first because most of the strategically interesting change is now happening at the catalog boundary.

Iceberg specifies table metadata and an open catalog interface, but Horizon and Unity add governance models, credential vending, federation, discovery, policies, audit capabilities, and product-specific control planes. Databricks, for example, now uses Unity Catalog’s Iceberg REST interface for managed Iceberg read/write access, while Snowflake uses Polaris inside Horizon to expose Snowflake-managed Iceberg tables to external engines.

Conflating format and catalog produces bad platform comparisons.

“Both vendors support Iceberg” tells you almost nothing about whether both allow the same external writers, preserve the same governance controls, vend temporary credentials, support v3, expose foreign tables, or let another engine commit transactions.

The concrete interview skill is therefore not reciting that Iceberg provides open tables. It is being able to draw:

writer → catalog → metadata → object storage → reader

Then explain who owns each arrow, which API crosses it, and where proprietary behavior still exists.

Certifications and Portfolio Signals Worth Having

There is no major vendor-neutral, dedicated Apache Iceberg certification that I could verify from the primary certification catalogs I checked. Apache Iceberg’s official site documents the project and specification but does not advertise an ASF Iceberg certification program, while Databricks’ certification catalog focuses on broader platform and engineering credentials.

However, saying “no certification tests catalog-layer knowledge” would be too strong.

Databricks’ Data Engineer Professional certification explicitly includes Unity Catalog among its platform competencies. It is not an Iceberg-specific credential, but catalog knowledge is clearly inside at least one mainstream data-engineering certification.

Signal

What it actually demonstrates

General data-engineering certification

Platform breadth and structured knowledge

Databricks professional certification

Includes Unity Catalog and production Databricks engineering

Dedicated Iceberg course/badge

Training exposure; not equivalent to an industry certification

Local Iceberg REST catalog lab

You can configure clients, catalogs, storage, and metadata

Format-migration project

You understand compatibility and migration mechanics

Cross-engine Spark/Trino/PyIceberg demo

You can prove interoperability rather than repeat vendor claims

For this specific topic, I would value a good portfolio repository more than an Iceberg multiple-choice badge.

Build a small dataset, register it through an Iceberg REST catalog, write with one engine, read with another, introduce a schema change, inspect snapshots, test an unsupported feature intentionally, and document what breaks. Then explain how the experiment would differ under Horizon Catalog, Unity Catalog, or S3 Tables.

A second strong project is a Delta-to-Iceberg interoperability lab rather than a simplistic “which format wins?” benchmark.

For example, demonstrate a Delta table exposed to Iceberg clients through UniForm, then compare that with a native managed Iceberg table. Databricks documents that UniForm generates Iceberg metadata without rewriting the shared Parquet files, so the exercise exposes the distinction between physical data, table metadata, and catalog access particularly well.

That is exactly the kind of project that gives an interviewer something concrete to probe.

What This Means for Data Engineer Salaries and Demand

This is not the place to rebuild a salary guide. Refonte Learning already publishes the full data engineer salary breakdown, including level-by-level compensation tables, so duplicating those figures here would create the content overlap this article is designed to avoid.

The more useful hiring question is whether table-format and catalog vocabulary is appearing in actual roles.

A current sample says yes. Cisco has advertised data-engineering roles explicitly naming Apache Iceberg alongside Spark, Kafka, Flink, and Trino; Barclays lists Unity Catalog in a Data Engineer role; UNSW asks for hands-on Unity Catalog expertise; and Hitachi has listed experience with Iceberg, Delta Lake, or Hudi for a senior Python data-engineering role.

Current hiring signal

Example

Apache Iceberg named directly

Cisco data-engineering roles

Unity Catalog named directly

Barclays, UNSW

Modern table-format experience accepted across formats

Hitachi: Iceberg, Delta Lake, or Hudi

Catalog/governance as a “plus”

SEP roles naming Unity Catalog

Iceberg within migration/lakehouse work

Hitachi roles using Iceberg in S3-oriented architectures

That evidence proves current demand, not a statistical increase over prior years. Without a longitudinal job-posting dataset, I would not claim that the percentage of roles mentioning Iceberg has risen by a specific amount.

Still, the shape of the requirements in current data-engineering roles is informative. Employers are not asking only whether a candidate “knows data lakes”; they are naming the governance and table layers used by the implementation.

That favors engineers who can discuss architecture trade-offs instead of presenting a tool checklist.

A senior candidate who can explain why an Iceberg REST endpoint changes migration options, why catalog federation differs from copying a table, or why a v3 writer can outrun reader compatibility is showing a more valuable skill than memorizing which company originally created which format.

Common Mistakes Teams Make Choosing a Table Format

Table-format decisions usually fail for reasons that have little to do with the headline comparison matrix.

Teams overestimate permanence, ignore the catalog, assume “open” means identical implementations, or benchmark read performance while overlooking commit behavior and maintenance. Those mistakes become more visible as how big data is driving innovation across data-driven organizations pushes more batch, streaming, AI, and governance workloads onto the same underlying data estate.

Mistake

Better question

“We chose Iceberg, so we solved lock-in”

Which proprietary catalog, policy, identity, optimizer, and storage dependencies remain?

“Delta or Iceberg is a permanent decision”

What would actually have to move to change formats or expose another representation?

“Supports Iceberg” equals interoperability

Which Iceberg spec version, REST operations, engines, and write paths are supported?

“Open format means all readers work”

Does every required engine support the features your writers emit?

“Zero ETL means no pipeline logic”

Which infrastructure job disappeared, and where do business transformations now run?

Treating Table-Format Choice as Permanent

The first mistake is assuming that choosing Delta Lake or Iceberg today commits the organization forever.

Databricks is a direct counterexample. UniForm can expose Delta-backed Parquet data to Iceberg readers, while Unity Catalog now also manages native Iceberg tables; Snowflake and Databricks have additionally been expanding catalog interoperability in 2026 rather than requiring every interaction to use a single proprietary table representation.

That does not make migration free.

Stored procedures, governance rules, clustering behavior, streaming semantics, engine-specific features, identity integrations, and operational automation can create far more lock-in than the files themselves.

So evaluate format choice against current workload needs. Do not select a technically inferior fit purely because you fear a migration that newer interoperability layers may make less disruptive.

Ignoring the Catalog Layer Entirely

The second mistake is treating the catalog as a dropdown you configure after selecting the table format.

Snowflake’s current Open Catalog/Horizon transition is exactly why that is unsafe. The table files can remain valid Iceberg while the supported route for discovering, governing, authenticating, and exposing those tables changes around them.

Ask catalog questions independently:

  • Who owns the authoritative metadata pointer?

  • How does the catalog coordinate concurrent commits?

  • Can external engines read?

  • Can they write?

  • Does the catalog vend scoped credentials?

  • Which governance policies survive external access?

  • Can you federate another catalog instead of migrating it?

  • What happens if the managed catalog product changes commercial or lifecycle direction?

Those questions will often decide the architecture before a benchmark of Delta versus Iceberg scan speed does.

Self-Study vs. a Structured Data Engineering Program: An Honest Comparison

You do not need a formal program to learn this material.

A strong engineer can build Spark pipelines, run Iceberg locally, study the specification, and test REST catalog behavior entirely through documentation and hands-on projects. The real trade-off is structure: self-study gives maximum flexibility, while a program gives a defined sequence, workload, mentorship, and evidence of completion.

The current Refonte Learning Data Engineering Program page verifies a three-month format, 12–14 hours per week, modules covering data-engineering foundations, ETL/data warehousing, and Big Data Technologies and Data Pipeline Design, with Hadoop and Spark named in the curriculum.

Factor

Self-study

Structured Data Engineering Program

Time to a working Spark pipeline

Varies with prior experience and study consistency

Spark sits inside the program’s Big Data Technologies and Data Pipeline Design module

ETL and warehouse depth

Depends on chosen projects

Dedicated Data Warehousing and ETL Processes module

Governance/compliance practice

Easy to skip unless deliberately planned

Listed among program competencies

Portfolio proof

Personal repositories and projects

Program page promises concrete projects and real-world experience; it does not specify a formal capstone-equivalent deliverable

Formal completion proof

None unless paired with another credential

Training Certificate + Certificate of Internship

Duration to job-ready fundamentals

Varies with prior experience, practice, and project scope

Program itself lasts 3 months; that does not guarantee job readiness

That last row needs the caveat.

A three-month program duration is an objective curriculum fact. “Job-ready in three months” would be a different claim, influenced by a learner’s starting skills, project quality, interview preparation, local market, and practice outside scheduled coursework; I would not infer it from the duration alone.

The same honesty applies to Iceberg.

The current Data Engineering Program page does not name Apache Iceberg, Delta Lake, Hudi, Airflow, Dagster, Prefect, Snowflake, or Databricks in its three-module curriculum. It explicitly names Hadoop and Spark in the big-data/pipeline module.

That means the defensible connection to this article is architectural foundation, not product-specific training.

Before debating Iceberg versus Delta, you need to understand why Spark reads distributed files, where transformations execute, what a pipeline does between streaming/batch ingestion and analytical storage, and why metadata coordination exists in the first place.

That foundation turns a format decision from a coin flip into engineering.

The Refonte Learning Data Engineering Program

The Refonte Learning Data Engineering Program is a three-month online training and virtual-internship program with a stated commitment of 12–14 hours per week. The current program page lists Data Engineer as the career outcome.

Its value in the context of open table formats is foundational.

It does not promise to teach Iceberg or Delta Lake. Instead, its curriculum covers the layers underneath those decisions: ETL, warehousing, big-data processing, Hadoop, Spark, streaming and batch ingestion competencies, scalable pipeline design, governance, storage provisioning, encryption, and security requirements.

The confirmed three-module curriculum is:

  • Introduction to Data Engineering: the role of data engineering and its place in the data lifecycle.

  • Data Warehousing and ETL Processes: warehousing, extract-transform-load concepts, and implementation.

  • Big Data Technologies and Data Pipeline Design: Hadoop, Spark, and the design of robust, scalable data pipelines.

That third module is the closest connection to the architecture discussed in this article. Understanding Spark as a distributed reader/writer and understanding how a batch or streaming pipeline lands and transforms data are prerequisites for evaluating what a table format and catalog are solving.

Program detail

Verified current information

Duration

3 months

Weekly commitment

12–14 hours/week

Format

Online / virtual-internship structure

Named big-data technologies

Apache Hadoop and Apache Spark

Curriculum modules

Introduction to Data Engineering; Data Warehousing and ETL Processes; Big Data Technologies and Data Pipeline Design

Mentor

PhD Matthias Schmidt, Department of Data Engineering

Mentor experience

16 years, including regression analysis, algorithmic design, financial econometrics, and big-data solutions for banking/financial services

Certificates

Training Certificate + Certificate of Internship

Top-performer recognition

Letter of Recommendation and Certificate of Appreciation may be awarded

One-time fee

$300

Installments

$204 + $98, totaling $302

Career result listed

Data Engineer

The program page identifies PhD Matthias Schmidt as its Department of Data Engineering mentor and describes a 16-year background spanning computer science, regression analysis, algorithmic design, financial econometrics, quantitative risk forecasting, and big-data solutions for banking and financial services.

The competencies listed on the live page include Big Data Analytics, provisioning data-storage services, encryption techniques, governance and compliance controls, real-time processing, data pipelining, visualization, ingesting streaming and batch data, transformations, and implementing security requirements.

On completion, the page says participants receive a Training Certificate and Certificate of Internship. It says outstanding performers may additionally receive a Letter of Recommendation and Certificate of Appreciation.

The program page currently markets a “$100.0K+ Starting” figure and “94K+ Jobs Annually” for Data Engineering. Those are Refonte Learning’s own marketing figures on the program page, not independent labor-market statistics verified in this article, so they should be treated as such.

The one-time enrollment price is $300. The two listed installment amounts are $204 and $98, which sum to $302 rather than $300.

The live page states the admission requirement as being engaged in or working toward a bachelor’s or higher-level degree.

For the pipeline-architecture foundation behind format and catalog decisions, see the Refonte Learning Data Engineering Program.

FAQ: People Also Ask

What changed with open table formats in 2026?

The strongest 2026 changes happened around catalogs, interoperability, and production implementation. Snowflake now directs first-time Open Catalog users to Horizon Catalog and made Horizon-based external Iceberg access generally available; Databricks made Managed Iceberg, Foreign Iceberg, and Iceberg v3 generally available in Unity Catalog on May 28, 2026.

AWS S3 Tables and Confluent Tableflow are also part of the consolidation, but their initial Iceberg moves predate 2026: S3 Tables launched in December 2024 and gained Iceberg REST-compatible management APIs in March 2025, while Tableflow for Iceberg became GA in Confluent Cloud in March 2025.

Is Snowflake's Open Catalog still available?

Yes for existing customers, but not for a customer creating its first Open Catalog account. Snowflake’s current documentation says first-time customers should use Snowflake Horizon Catalog for Iceberg and multi-engine interoperability, while eligible existing Open Catalog customers can continue operating and can create additional accounts.

What does Iceberg v3 actually add?

Iceberg v3 adds nanosecond timestamp types; unknown, variant, geometry, and geography types; default column values; multi-argument partitioning and sorting transforms; row lineage; binary deletion vectors; and table encryption-key metadata. The Apache Iceberg project currently says v3 is complete and adopted by the community, while v4 remains under development.

Does Databricks support Iceberg, or only Delta Lake?

Databricks now supports both. Unity Catalog exposes an Iceberg REST Catalog implementation, Managed Iceberg supports external read/write access, and Delta tables can be configured for Iceberg reads through UniForm, which generates Iceberg metadata without rewriting the underlying Parquet data.

Unity Catalog’s open-sourcing under Apache 2.0 through LF AI & Data occurred on June 13, 2024, so that event should not be presented as a 2026 announcement.

Do I need to migrate my table format because of these 2026 changes?

No. These changes improve interoperability and add managed infrastructure; they do not create a universal forced migration from Delta Lake, Hudi, earlier Iceberg versions, or conventional Iceberg deployments.

Evaluate a migration when you can identify a measurable benefit such as eliminating a custom materialization pipeline, improving external-engine access, consolidating governance, reducing table-maintenance work, or removing a catalog limitation. Also verify reader/writer support before upgrading Iceberg format versions because implementation can lag the specification.

Does the Data Engineering program teach Iceberg or Delta Lake specifically?

No. The current Refonte Learning Data Engineering Program page names Hadoop and Spark in its Big Data Technologies and Data Pipeline Design module; it does not name Iceberg, Delta Lake, Hudi, Snowflake, or Databricks in the confirmed three-module curriculum.

The program’s relevance here is foundational: it teaches data-engineering, ETL, warehousing, big-data, and pipeline-design concepts that you need before evaluating where a table format or catalog fits in an actual architecture.

The 2026 picture is clearer once you separate genuine standards movement from product marketing:

  • 2026 is a real consolidation moment, especially because Snowflake moved its strategic Iceberg interoperability path toward Horizon Catalog while Databricks made managed Iceberg and v3 support GA. AWS and Confluent reinforce that shift, although their foundational Iceberg products originated before 2026.

  • The catalog layer is where much of the important action now happens. Horizon Catalog, Unity Catalog, Apache Polaris, REST endpoints, credential vending, federation, and external-engine write paths matter as much as the files in object storage.

  • Iceberg v3 is a production-maturity milestone, adding row identity, deletion vectors, richer types, defaults, transforms, and encryption-key metadata while the project now lists v3 as complete and adopted.

  • None of this justifies a reactive migration. Test your engines, catalogs, permissions, format versions, operational costs, and failure modes first; an open specification reduces friction only where your actual implementations honor it.

For the pipeline-architecture foundation that makes those table-format and catalog decisions meaningful rather than guesswork, the Refonte Learning Data Engineering Program is the structured starting point described above.