Handling schema inference and nested JSON issues from web traffic logs in S3

Hello everyone, I am currently setting up an analytical pipeline in Dremio to analyze user session metrics and API payload dumps collected from this website. We sync the raw incoming event logs into an AWS S3 bucket in JSON format, but whenever I try to promote the parent folder into a physical dataset, Dremio fails to infer the schema properly due to mixed data types and deeply nested structures in the event objects.

Has anyone encountered similar issues where semi-structured web telemetry causes metadata sync delays or unexpected type mismatches during query execution? I would really appreciate advice on whether it is safer to flatten these dynamic fields prior to ingestion or if relying on virtual datasets with Arrow-optimized JSON parsing functions can handle this variability without degrading reflection performance.