# 1 of the 4 worker node dead

**URL:** <https://community.dremio.com/t/1-of-the-4-worker-node-dead/11332>\
**Category:** Uncategorized\
**Created:** [January 3, 2024, 3:04am UTC](https://community.dremio.com/t/1-of-the-4-worker-node-dead/11332 "2024-01-03T03:04:53Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![crazyisjen](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/crazyisjen/32/4552_2.png) [@crazyisjen](https://community.dremio.com/u/crazyisjen)\
**Post date:** [January 3, 2024, 3:04am UTC](https://community.dremio.com/t/1-of-the-4-worker-node-dead/11332/1 "2024-01-03T03:04:53Z")

</div>

> Exceeded timeout (30000) while waiting after sending work fragments to remote nodes. Sent 3 and only heard response back from 2 nodes

I have experience above error frequently, guess probably occur when running complex query concurrently, I have limit the concurrent query from 8 to 2 but still a node will dead.

Could you pls advise how to limit the memory allocation to each query as to avoid any work node fails and cause impact to other queries. Thanks a lot~

---

<div class="post-metadata">

**Author:** ![bogdan.coman](https://avatars.discourse-cdn.com/v4/letter/b/f6c823/32.png) [@bogdan.coman](https://community.dremio.com/u/bogdan.coman)\
**Post date:** [January 4, 2024, 3:01am UTC](https://community.dremio.com/t/1-of-the-4-worker-node-dead/11332/2 "2024-01-04T03:01:04Z")

</div>

Hi @crazyisjen,

In the community version you have the Query Memory Control settings you can use (not per query, but for small and large queries): [Queue Control | Dremio Documentation](https://docs.dremio.com/current/admin/workloads/job-queues/#query-memory-control)

Thanks, Bogdan

---

<div class="post-metadata">

**Author:** ![crazyisjen](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/crazyisjen/32/4552_2.png) [@crazyisjen](https://community.dremio.com/u/crazyisjen)\
**Post date:** [January 10, 2024, 6:45am UTC](https://community.dremio.com/t/1-of-the-4-worker-node-dead/11332/3 "2024-01-10T06:45:01Z")

</div>

thanks a lot  
Would like to ask more abt Query Threshold

I saw default 30000000 is enabled , what is the units of this figures? And is it a ceiling for each query to consume the memory?  
And I have check one of the large successful completed query which took 3401436399 in query cost, any reason it is not stopped if exceed the ceiling?

thanks a lot for help 😃

---

<div class="post-metadata">

**Author:** ![balaji.ramaswamy](https://sea2.discourse-cdn.com/flex020/user_avatar/community.dremio.com/balaji.ramaswamy/32/843_2.png) [@balaji.ramaswamy](https://community.dremio.com/u/balaji.ramaswamy)\
**Post date:** [January 14, 2024, 12:44am UTC](https://community.dremio.com/t/1-of-the-4-worker-node-dead/11332/4 "2024-01-14T00:44:23Z")

</div>

@crazyisjen The error you are getting means the executor did not respond, this could be due to 2 reasons

- Full GC pause
- RPC’s between Dremio nodes taking too much time

When this happens next time, send here the below

- The job profile of the job that failed
- server.log from the executor mentioned in the error that did not respond
- GC logs from the executor mentioned in the error that did not respond

IF you are K8’s then use below parameters in values.yaml to first move GC logging to a PVC

Open `values.yaml` and add the following under the appropriate section; executor and/or coordinator:

extraStartParams: \>-  
-Xloggc:/opt/dremio/data/gc-%t-%p.log  
-XX:+UseGCLogFileRotation  
-XX:NumberOfGCLogFiles=5  
-XX:GCLogFileSize=4000k  
-XX:+PrintGCDetails  
-XX:+PrintGCTimeStamps  
-XX:+PrintGCDateStamps  
-XX:+PrintClassHistogramBeforeFullGC  
-XX:+PrintClassHistogramAfterFullGC  
-XX:+HeapDumpOnOutOfMemoryError  
-XX:HeapDumpPath=/opt/dremio/data  
-XX:+UseG1GC  
-XX:G1HeapRegionSize=32M  
-XX:MaxGCPauseMillis=500  
-XX:InitiatingHeapOccupancyPercent=25  
-XX:+PrintAdaptiveSizePolicy  
-XX:+PrintReferenceGC  
-XX:ErrorFile=/opt/dremio/data/hs\_err\_pid%p.log
