# Reading Dremio's parquet files from python

**URL:** <https://community.dremio.com/t/reading-dremios-parquet-files-from-python/3702>\
**Category:** Uncategorized\
**Created:** [July 18, 2019, 9:18am UTC](https://community.dremio.com/t/reading-dremios-parquet-files-from-python/3702 "2019-07-18T09:18:59Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![romain](https://avatars.discourse-cdn.com/v4/letter/r/c2a13f/32.png) [@romain](https://community.dremio.com/u/romain)\
**Post date:** [July 18, 2019, 9:18am UTC](https://community.dremio.com/t/reading-dremios-parquet-files-from-python/3702/1 "2019-07-18T09:18:59Z")

</div>

Hi

Over the last year, I’ve been successfully generating parquet from from python and issuing queries on them using Dremio, all this works perfectly.

However some of these tables are large denormalized files and take forever to create in python. I was considering delegating the file creation to Dremio, and was hoping to be able to read them back both from Dremio and python.

Hence I’ve created a table from Dremio using  
CREATE TABLE “TESTS\_CTAS”.“CTAS\_V\_COUNTRY”  
AS SELECT \* FROM PATH\_TO\_VDS.V\_COUNTRY

This generates a folder CTAS\_V\_COUNTRY on the filesystem, which contains two files : 0\_0\_0.parquet and 0\_0\_0.parquet.crc

I was hoping to read back the file using pyarrow. I can access the schema and the metadata using  
f = pq.ParquetFile(os.path.join(ctas\_path, ‘CTAS\_V\_COUNTRY’, ‘0\_0\_0.parquet’))  
f.metadata, f.schema

However, reading the file  
pq.read\_table(os.path.join(ctas\_path, ‘CTAS\_V\_COUNTRY’, ‘0\_0\_0.parquet’))  
fails and yields the following error  
ArrowIOError: Couldn’t deserialize thrift: TProtocolException: Invalid data  
Deserializing page header failed.

Of course, I could read the file using pyodbc, but this is painfully slow. I was hoping that the the newer Flight api could help, but I couldn’t make what you refer to in your blog work : [https://www.dremio.com/is-time-to-replace-odbc-jdbc/](https://www.dremio.com/is-time-to-replace-odbc-jdbc/)

import pyarrow.flight as flt  
c = flt.FlightClient.connect(“localhost”, 47470)

yields  
TypeError: expected bytes, int found

I’m not too sure how to pass the port, and which port I should give, anyway I have guessed wrong since using  
c = flt.FlightClient.connect(“localhost”, (47470).to\_bytes(2, byteorder=‘big’))  
later fails on  
fi = c.get\_flight\_info(fd)  
-\>  
ArrowIOError: gRPC failed with error code 14 and message: DNS resolution failed

What would be your recommendation to achieve this objective of creating parquet using Dremio and reading it back fast in python (both Dremio and my python code run on the same server) ?

Thanks for your help,  
Romain

PS : I use Dremio 3.1.11-201904261857420193-c674472 and pyarrow 0.14.0

---

<div class="post-metadata">

**Author:** ![dfleckinger](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/dfleckinger/32/370_2.png) [@dfleckinger](https://community.dremio.com/u/dfleckinger)\
**Post date:** [July 29, 2019, 1:37pm UTC](https://community.dremio.com/t/reading-dremios-parquet-files-from-python/3702/2 "2019-07-29T13:37:33Z")

</div>

You should try

c = flight.FlightClient.connect(‘grpc+tcp://localhost:47470’)

> **[dremio-hub/dremio-flight-connector](https://github.com/dremio-hub/dremio-flight-connector)**
>
> Dremio Flight connector. Access Dremio using Arrow flight - dremio-hub/dremio-flight-connector
