# Desk log FLF-PDA-20260822

**Status:** First-party InfiniSynapse desk log (sanitized composite; not a customer extract)  
**Page:** https://infinisynapse.com/en/blog/parquet-data-analysis  
**Run ID:** `FLF-PDA-20260822`  
**Date:** 2026-08-22 (Saturday)  
**Operator:** InfiniSynapse Data Team  
**Attestor:** William Zhu, InfiniSynapse cofounder ([GitHub @allwefantasy](https://github.com/allwefantasy))  
**Contact for contradictions:** zhuhl@infinisynapse.com

## What this file is

A downloadable record of converting a weekly folder to CSV first versus asking the folder in place. It is **not** a named-logo customer case, a vendor bake-off, or an Apache / BigQuery / Azure experiment.

## Four-step method (reproducible)

1. Register the twelve weekly partitions, format, and owner. Sanitize first.
2. Profile schema, partitions, nulls, and a row-count check. Note week-9 drift.
3. Ask one goal: return rate by SKU for the last six weeks, with orders in the denominator.
4. Open sampling versus full scan and denominator counts. Re-ask the same goal from a clean seat.

## Source and goal

| Field | Value |
|---|---|
| Sources | Twelve weekly parquet partitions, about 4.2 million rows |
| Standing goal | Return rate by SKU for the last six weeks, excluding one-week-only SKUs |
| Contrast | Convert to CSV first vs ask the folder |

## Results

| Retrieval state | Week-9 drift profiled | Denominator counts opened | Same-day re-ask possible |
|---|---|---|---|
| Convert to CSV first | 0 | 0 | 0 |
| Ask the folder | 1 | 1 | 1 |

Wall-clock for the successful folder rerun: 10 minutes (warehouse time excluded). The clock started when the operator opened the standing goal and ended when both folders sat side by side with the week-9 schema note and the denominator counts open.

## What you may cite

- Artifact counts 0/0/0 → 1/1/1, 12 weekly partitions + ~4.2M rows on this run, ~10 min wall-clock, run ID

## What you may not claim

- Customer uplift %, a 40% faster scan, official EEAT score, named-logo case, or that Apache / BigQuery / Azure / Gartner / WAIC scored this run

## Independent context (not this run)

- [Apache Parquet documentation](https://parquet.apache.org/docs/) (retrieved 2026-08-29)
- [Apache parquet-format](https://github.com/apache/parquet-format) (retrieved 2026-08-29)
- [Wikipedia Apache Parquet](https://en.wikipedia.org/wiki/Apache_Parquet) (retrieved 2026-08-29)
- [Apache Arrow documentation](https://arrow.apache.org/docs/) (retrieved 2026-08-29)
- [DuckDB Parquet guide](https://duckdb.org/docs/stable/data/parquet/overview.html) (retrieved 2026-08-29)
- [ISO/IEC 9075](https://www.iso.org/standard/76583.html) (retrieved 2026-08-29)
- [RFC 4180 CSV](https://www.rfc-editor.org/rfc/rfc4180) (retrieved 2026-08-29)
- [pandas documentation](https://pandas.pydata.org/docs/) (retrieved 2026-08-29)
- [Apache Spark documentation](https://spark.apache.org/docs/latest/) (retrieved 2026-08-29)
- [PostgreSQL documentation](https://www.postgresql.org/docs/) (retrieved 2026-08-29)
- [Google BigQuery documentation](https://cloud.google.com/bigquery/docs) (retrieved 2026-08-29)
- [Microsoft Azure data architecture guide](https://learn.microsoft.com/en-us/azure/architecture/data-guide/) (retrieved 2026-08-29)
- First-party homepage recognition only: [2026 WAIC Future Tech OPC Excellence Award](https://infinisynapse.com/#recognition) (self-described; not independently verified here)

## Machine-readable aggregate

- [aggregate-FLF-PDA-20260822.csv](./aggregate-FLF-PDA-20260822.csv)
- [verify-FLF-PDA-20260822.py](./verify-FLF-PDA-20260822.py)
