# Number of Threads

**URL:** <https://community.dremio.com/t/number-of-threads/7198>\
**Category:** Uncategorized\
**Created:** [March 24, 2021, 6:17pm UTC](https://community.dremio.com/t/number-of-threads/7198 "2021-03-24T18:17:00Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![tylercmp](https://avatars.discourse-cdn.com/v4/letter/t/59ef9b/32.png) [@tylercmp](https://community.dremio.com/u/tylercmp)\
**Post date:** [March 24, 2021, 6:17pm UTC](https://community.dremio.com/t/number-of-threads/7198/1 "2021-03-24T18:17:00Z")

</div>

Hi,

I have Dremio deployed to an EKS cluster and have a couple question:

- Is there a way to configure the number of threads used?
- Any best practices on this configuration?
- What is the default number of threads used?

Thanks

---

<div class="post-metadata">

**Author:** ![balaji.ramaswamy](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/balaji.ramaswamy/32/843_2.png) [@balaji.ramaswamy](https://community.dremio.com/u/balaji.ramaswamy)\
**Post date:** [March 25, 2021, 7:48am UTC](https://community.dremio.com/t/number-of-threads/7198/2 "2021-03-25T07:48:27Z")

</div>

@tylercmp

When you say threads, is it for the scans? # of threads depends on 3 factors

- Number of input splits
- Number of cores per executor (all executors should have same number of cores)
- Estimated number of rows on the scan

We should leave it to Dremio to decide, are you facing a performance issue? If yes, kindly share the query profile

> **[How To Share A Query Profile](https://www.dremio.com/tutorials/share-query-profile-dremio/)**
>
> Tutorial explaining how to share a query profile in Dremio.

---

<div class="post-metadata">

**Author:** ![tylercmp](https://avatars.discourse-cdn.com/v4/letter/t/59ef9b/32.png) [@tylercmp](https://community.dremio.com/u/tylercmp)\
**Post date:** [March 25, 2021, 7:34pm UTC](https://community.dremio.com/t/number-of-threads/7198/3 "2021-03-25T19:34:09Z")

</div>

My cluster is set up with 2 coordinators and 3 executors. My use case is small, low latency queries at a high frequency. At a moderate load I was seeing good performance and low resource utilization (CPU and Memory). After increasing to a higher load, I’m seeing degraded performance, but poor resource utilization.

**Moderate Load Query Profiles:**  
[a5476b14-1dff-46be-ad3a-b1c1f56ba912.zip](https://community.dremio.com/uploads/short-url/fR8Hj0iBHitWfn9IHRMghIt3hU6.zip) (14.0 KB)

**High Load Query Profiles:**  
[07bd3b77-cb6f-4168-a7f2-0d398426cbf5.zip](https://community.dremio.com/uploads/short-url/5Xa80uVSL94RpgI0wxNHlclr8tJ.zip) (28.1 KB)

**High Load CPU/Memory:**

 ![Screen Shot 2021-03-25 at 2.20.10 PM](https://us1.discourse-cdn.com/flex020/uploads/dremio/original/2X/1/1c24baad23132d1298183d7cfd33c9be0154526f.png)

---

<div class="post-metadata">

**Author:** ![tylercmp](https://avatars.discourse-cdn.com/v4/letter/t/59ef9b/32.png) [@tylercmp](https://community.dremio.com/u/tylercmp)\
**Post date:** [March 25, 2021, 7:35pm UTC](https://community.dremio.com/t/number-of-threads/7198/4 "2021-03-25T19:35:46Z")

</div>

Additional Query Profiles:

**Moderate Load Query Profiles:**  
[544ace0c-e972-47ae-b691-a9e32d11106d.zip](https://community.dremio.com/uploads/short-url/hiQiEULOOZb1sYsoWbgFjV129V4.zip) (12.1 KB)

**High Load Query Profiles:**  
[084ce920-1649-4333-9fd2-0bd943b2c7fe.zip](https://community.dremio.com/uploads/short-url/bDnsqZ1piSKlpdWULyILEzt34Qn.zip) (13.9 KB)

---

<div class="post-metadata">

**Author:** ![balaji.ramaswamy](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/balaji.ramaswamy/32/843_2.png) [@balaji.ramaswamy](https://community.dremio.com/u/balaji.ramaswamy)\
**Post date:** [March 26, 2021, 5:55am UTC](https://community.dremio.com/t/number-of-threads/7198/5 "2021-03-26T05:55:58Z")

</div>

@tylercmp

The queries are completing in 135ms, what is your expectation?

---

<div class="post-metadata">

**Author:** ![alex.shi](https://avatars.discourse-cdn.com/v4/letter/a/ea666f/32.png) [@alex.shi](https://community.dremio.com/u/alex.shi)\
**Post date:** [December 29, 2021, 9:00am UTC](https://community.dremio.com/t/number-of-threads/7198/6 "2021-12-29T09:00:36Z")

</div>

> [@balaji.ramaswamy](#):
>
> Estimated number of rows on the scan

[2min.zip](https://community.dremio.com/uploads/short-url/qJRwtqbKgVCp5VABbzjESujpGYb.zip) (2.5 MB)  
@balaji.ramaswamy Can you help me optimize this query? I don’t understand the sqlprofile, Please! TKS!

---

<div class="post-metadata">

**Author:** ![balaji.ramaswamy](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/balaji.ramaswamy/32/843_2.png) [@balaji.ramaswamy](https://community.dremio.com/u/balaji.ramaswamy)\
**Post date:** [January 3, 2022, 5:38am UTC](https://community.dremio.com/t/number-of-threads/7198/7 "2022-01-03T05:38:01Z")

</div>

@alex.shi Here is where the time is spent

01-xx-01 ARROW\_WRITER 8s wait time on IO, are you writing results to local disk or a distributed storage like S3/HDFS? When writing to a distributed storage, this wait time is sometimes expected

01-xx-03 PROJECT has a SETUP time of 27s, this is due to the Gandiva expression getting evaluated, if you re run the query, it should go away. If it is still there then we can look at increasing the cache size but for that we need to make sure we have enough memory from the OS available

02-xx-00 HASH\_PARTITION\_SENDER and 02-xx-01 PROJECT ~ 34s on SETUP, the PROJECT is again due to the above Gandiva expression build time

There a few FILTER (like 06-xx-03 FILTER) operators that have taken ~ 5s each on SEUP for the same Gandiva setup time,

Phases 1,2 and 14 are high on parallelism and this cause some CPU contention as shown under sleep time, are there other queries running at the same time?

---

<div class="post-metadata">

**Author:** ![alex.shi](https://avatars.discourse-cdn.com/v4/letter/a/ea666f/32.png) [@alex.shi](https://community.dremio.com/u/alex.shi)\
**Post date:** [January 4, 2022, 3:07am UTC](https://community.dremio.com/t/number-of-threads/7198/8 "2022-01-04T03:07:44Z")

</div>

Thank you for your feedback, and NOW I will answer your questions:

1. I write results to HDFS.

2. I have one coordination node and nineteen execution nodes.  
The Server physical memory of the coordination node is 125GB, and I set DREMIO\_MAX\_PERMGEN\_MEMORY\_SIZE\_MB=110000, DREMIO\_MAX\_HEAP\_MEMORY\_SIZE\_MB= 80000.  
The Server physical memory of the execution node is 256GB. and I set DREMIO\_MAX\_PERMGEN\_MEMORY\_SIZE\_MB=120000, DREMIO\_MAX\_HEAP\_MEMORY\_SIZE\_MB= 60000.  
I disabled swap, but when I run the top command to check the memory, I found that the physical memory usage was small and the virtual memory usage was large. Physical memory usage is low

3. There is only one query at a time. I don’t know why is the concurrency high.

I look forward to your answers!@balaji.ramaswamy

---

<div class="post-metadata">

**Author:** ![alex.shi](https://avatars.discourse-cdn.com/v4/letter/a/ea666f/32.png) [@alex.shi](https://community.dremio.com/u/alex.shi)\
**Post date:** [January 12, 2022, 7:35am UTC](https://community.dremio.com/t/number-of-threads/7198/9 "2022-01-12T07:35:26Z")

</div>

I have to say that the lack of feedback/very slow response from Dremio.  
I found the sqlprofile really hard to understand and not well documented.

---

<div class="post-metadata">

**Author:** ![balaji.ramaswamy](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/balaji.ramaswamy/32/843_2.png) [@balaji.ramaswamy](https://community.dremio.com/u/balaji.ramaswamy)\
**Post date:** [January 13, 2022, 7:57am UTC](https://community.dremio.com/t/number-of-threads/7198/10 "2022-01-13T07:57:51Z")

</div>

@alex.shi

Apologies for the delay, was there a reason to have such a high heap memory? Dremio execution all happens on direct memory and hence on the executor you can set the below values

comment out below 2,  
#DREMIO\_MAX\_MEMORY\_SIZE\_MB=  
#DREMIO\_MAX\_PERMGEN\_MEMORY\_SIZE\_MB=

enable below 2  
DREMIO\_MAX\_HEAP\_MEMORY\_SIZE\_MB= 8192  
DREMIO\_MAX\_DIRECT\_MEMORY\_SIZE\_MB=112640 # we can increase this later if required

The coordinator is heap intensive, so the below values,  
comment out below 2,  
#DREMIO\_MAX\_MEMORY\_SIZE\_MB=  
#DREMIO\_MAX\_PERMGEN\_MEMORY\_SIZE\_MB=

enable below 2  
DREMIO\_MAX\_HEAP\_MEMORY\_SIZE\_MB= 16384  
DREMIO\_MAX\_DIRECT\_MEMORY\_SIZE\_MB=16384

If the coordinator or executors runs out of heap or has long GC pauses, we can enable histograms to find out why and then take next steps

---

<div class="post-metadata">

**Author:** ![alex.shi](https://avatars.discourse-cdn.com/v4/letter/a/ea666f/32.png) [@alex.shi](https://community.dremio.com/u/alex.shi)\
**Post date:** [January 13, 2022, 9:16am UTC](https://community.dremio.com/t/number-of-threads/7198/11 "2022-01-13T09:16:22Z")

</div>

Thank you for your feedback！！  
I reconfigured according to this parameter, and the performance is no better.  
I’d like to ask you one more question，since HDFS input is slow, I have disabled the “enable local caching for hdfs” option to see if there is a way to specify caching to local disk instead of HDFS.Maybe that’ll give me a little bit of a performance boost.@balaji.ramaswamy

---

<div class="post-metadata">

**Author:** ![balaji.ramaswamy](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/balaji.ramaswamy/32/843_2.png) [@balaji.ramaswamy](https://community.dremio.com/u/balaji.ramaswamy)\
**Post date:** [January 18, 2022, 4:16pm UTC](https://community.dremio.com/t/number-of-threads/7198/12 "2022-01-18T16:16:51Z")

</div>

Caching should help with the wait times incurred
