# Dremio Iceberg JDBC catalog

**URL:** <https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614>\
**Category:** Uncategorized\
**Created:** [May 16, 2023, 7:11am UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614 "2023-05-16T07:11:52Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![txalaparta](https://avatars.discourse-cdn.com/v4/letter/t/919ad9/32.png) [@txalaparta](https://community.dremio.com/u/txalaparta)\
**Post date:** [May 16, 2023, 7:11am UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/1 "2023-05-16T07:11:52Z")

</div>

In my working scenario I have a client environment written in JAVA and setup a development architecture deploying [GitHub - tabular-io/docker-spark-iceberg](https://github.com/tabular-io/docker-spark-iceberg) docker images and S3 compatible local storage (minio). In this case it using a rest iceberg catalog which is backed by a sqlite database. I am able to create and manipulate the iceberg tables programatically in spark (javas, sql, python) but when I try to connect them to dremio and this is the error I get:  
_This folder does not contain a filesystem-based Iceberg table. If the table in this folder is managed via a catalog such as Hive, Glue, or Nessie, please use a data source configured for that catalog to connect to this table._

 ![image](https://us1.discourse-cdn.com/flex020/uploads/dremio/original/2X/c/c24c913e9b8ed883c47d31cfd6f4c1315e352f90.png)

On the other hand I have created successfully an Iceberg table from Dremio in the same Minio bucket.  
I am wondering if a jdbc catalog (I could use postgreSQL for example) would be recognised by dremio… Or should I install HIVE and connect it to minio?  
I would like to avoid the need for a spark environment if possible.  
Thanks  
Oskar

---

<div class="post-metadata">

**Author:** ![YuriyGavrilov](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/yuriygavrilov/32/4875_2.png) [@YuriyGavrilov](https://community.dremio.com/u/YuriyGavrilov)\
**Post date:** [May 16, 2023, 6:41pm UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/2 "2023-05-16T18:41:22Z")

</div>

same message, same problem.

---

<div class="post-metadata">

**Author:** ![balaji.ramaswamy](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/balaji.ramaswamy/32/843_2.png) [@balaji.ramaswamy](https://community.dremio.com/u/balaji.ramaswamy)\
**Post date:** [May 17, 2023, 1:26am UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/3 "2023-05-17T01:26:21Z")

</div>

@txalaparta When you created the Iceberg table, which catalog was set?

---

<div class="post-metadata">

**Author:** ![txalaparta](https://avatars.discourse-cdn.com/v4/letter/t/919ad9/32.png) [@txalaparta](https://community.dremio.com/u/txalaparta)\
**Post date:** [May 18, 2023, 5:41am UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/4 "2023-05-18T05:41:36Z")

</div>

I created the Iceberg table with a REST catalog built in tabulario/iceberg-rest docker image.

I also created another table in dremio and this saves data in minio. However cannot connect to it from JAVA API nor from PyIcberg. What is the type of the catalog created in dremio? According to dremio documentation at [Dremio](https://docs.dremio.com/software/data-formats/apache-iceberg/), mino should be a Hadoop Iceberg catalog right?  
Thanks  
Oskar

---

<div class="post-metadata">

**Author:** ![nicolas.guerra](https://avatars.discourse-cdn.com/v4/letter/n/8baadc/32.png) [@nicolas.guerra](https://community.dremio.com/u/nicolas.guerra)\
**Post date:** [May 18, 2023, 2:21pm UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/5 "2023-05-18T14:21:05Z")

</div>

we are using dremio with minio and apache iceberg, by the moment the best result are with the iceberg catalog on hadoop type in the same bucket, because on dremio you must configure the catalog not the bucket on minio.  
this is and examplo to create with spark the catalog on minio bucket,  
spark.conf.set(‘spark.sql.catalog.silver\_data’, ‘org.apache.iceberg.spark.SparkCatalog’)  
spark.conf.set(‘spark.sql.catalog.silver\_data.type’, ‘hadoop’)  
spark.conf.set(‘spark.sql.catalog.silver\_data.warehouse’, ‘s3a://warehouse-silver-data’)  
then the table metadata and data area written in this bucket, and dremio can read that bucket and you can take every folder as table and format as apache iceberg,  
we are wating for nessie integration…

---

<div class="post-metadata">

**Author:** ![txalaparta](https://avatars.discourse-cdn.com/v4/letter/t/919ad9/32.png) [@txalaparta](https://community.dremio.com/u/txalaparta)\
**Post date:** [May 19, 2023, 5:56am UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/6 "2023-05-19T05:56:14Z")

</div>

Thanks for the info Nicolas.  
It is very helpful.  
I will try to configure a hadoop type catalog,

---

<div class="post-metadata">

**Author:** ![YuriyGavrilov](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/yuriygavrilov/32/4875_2.png) [@YuriyGavrilov](https://community.dremio.com/u/YuriyGavrilov)\
**Post date:** [May 22, 2023, 2:53pm UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/7 "2023-05-22T14:53:17Z")

</div>

I used trino icberg connector to create catalog but it does not work. Just can’t read (same message) [Iceberg connector — Trino 418 Documentation](https://trino.io/docs/current/connector/iceberg.html)

---

<div class="post-metadata">

**Author:** ![YuriyGavrilov](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/yuriygavrilov/32/4875_2.png) [@YuriyGavrilov](https://community.dremio.com/u/YuriyGavrilov)\
**Post date:** [May 23, 2023, 8:26am UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/8 "2023-05-23T08:26:09Z")

</div>

Thanks @nicolas.guerra

Tryed [Docker, Spark, and Iceberg: The Fastest Way to Try Iceberg! • Tabular](https://tabular.io/blog/docker-spark-and-iceberg/) as rest catalog for the Trino ([Iceberg connector — Trino 418 Documentation](https://trino.io/docs/current/connector/iceberg.html)) Then create iceberg table in s3 path using Trino.  
But anyway Dremio can’t read this folder and i receive same message “This folder does not contain filesystem based Iceberg table…” also tryed different s3 provider but same…

---

<div class="post-metadata">

**Author:** ![txalaparta](https://avatars.discourse-cdn.com/v4/letter/t/919ad9/32.png) [@txalaparta](https://community.dremio.com/u/txalaparta)\
**Post date:** [May 23, 2023, 8:43am UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/9 "2023-05-23T08:43:52Z")

</div>

I´ve left this issue aside for a while and not planning to continue yet. In any case I think the solution is by implementing a **hadoop** or **hive** catalog instead of jdbc or rest. I am quite sure the last two will not work in dremio.  
I saw some documentation on how to configure hadoop to connect to Minio and also some github to create such a catalog in Trino. (Sorry but I don´t have the links right now)  
But yes, I seems that Spark (or Trino) is needed.

Good luck

---

<div class="post-metadata">

**Author:** ![nicolas.guerra](https://avatars.discourse-cdn.com/v4/letter/n/8baadc/32.png) [@nicolas.guerra](https://community.dremio.com/u/nicolas.guerra)\
**Post date:** [May 24, 2023, 7:44pm UTC](https://community.dremio.com/t/dremio-iceberg-jdbc-catalog/10614/10 "2023-05-24T19:44:47Z")

</div>

i think the best way to work with minio iceberg, its getting the nessie conector, its pretty useful.  
right now, we are using airflow-spark solution that can sabe data in raw format in iceberg tables using hadoop catalog, then with dremio we read all data formating the folder that contains data and metadata folders, for ddl task and datamanagement we are using jupyter notebooks.
