# Desk log FLF-CVP-20260822

**Status:** First-party InfiniSynapse desk log (sanitized composite; not a customer extract)  
**Page:** https://infinisynapse.com/en/blog/csv-vs-parquet-for-ai  
**Run ID:** `FLF-CVP-20260822`  
**Date:** 2026-08-22 (Saturday)  
**Operator:** InfiniSynapse Data Team  
**Attestor:** William Zhu, InfiniSynapse cofounder ([GitHub @allwefantasy](https://github.com/allwefantasy))  
**Contact for contradictions:** zhuhl@infinisynapse.com

## What this file is

A downloadable record of asking a CSV email first versus asking the typed Parquet of the same week. It is **not** a named-logo customer case, a vendor bake-off, or an ACM / arXiv / Science experiment.

## Four-step method (reproducible)

1. Register both paths, owner, and allowed use. Sanitize first.
2. Profile CSV quoting, encoding, and null tokens. Profile Parquet schema.
3. Ask one goal: return rate by SKU, denominator = orders, on one winner.
4. Open types and row counts. Name the winner. Re-ask the same goal from a clean seat.

## Source and goal

| Field | Value |
|---|---|
| Sources | One week as a 4.2-million-row Parquet and a CSV email of the same grain |
| Standing goal | Return rate by SKU, denominator = orders |
| Contrast | CSV email first vs ask the typed Parquet |

## Results

| Retrieval state | Orders typed int64 | Thousands-separator nulls caught | Winner named |
|---|---|---|---|
| CSV email first | 0 | 0 | 0 |
| Ask the typed Parquet | 1 | 1 | 1 |

Wall-clock for the successful typed rerun: 10 minutes (warehouse time excluded). The clock started when the operator opened the standing goal and ended when both files sat side by side with the type clash and the winner note open.

## What you may cite

- Artifact counts 0/0/0 → 1/1/1, one 4.2M-row week as CSV + Parquet on this run, ~10 min wall-clock, run ID

## What you may not claim

- Customer uplift %, a faster scan, official EEAT score, named-logo case, or that ACM / arXiv / Science / Gartner / WAIC scored this run

## Independent context (not this run)

- [RFC 4180 CSV](https://www.rfc-editor.org/rfc/rfc4180) (retrieved 2026-08-29)
- [Apache Parquet documentation](https://parquet.apache.org/docs/) (retrieved 2026-08-29)
- [Apache parquet-format](https://github.com/apache/parquet-format) (retrieved 2026-08-29)
- [Apache Arrow documentation](https://arrow.apache.org/docs/) (retrieved 2026-08-29)
- [DuckDB CSV guide](https://duckdb.org/docs/stable/data/csv/overview.html) (retrieved 2026-08-29)
- [DuckDB Parquet guide](https://duckdb.org/docs/stable/data/parquet/overview.html) (retrieved 2026-08-29)
- [W3C tabular data model](https://www.w3.org/TR/tabular-data-model/) (retrieved 2026-08-29)
- [ISO/IEC 9075](https://www.iso.org/standard/76583.html) (retrieved 2026-08-29)
- [Wikipedia Apache Parquet](https://en.wikipedia.org/wiki/Apache_Parquet) (retrieved 2026-08-29)
- [pandas I/O — CSV](https://pandas.pydata.org/docs/user_guide/io.html#io-read-csv-table) (retrieved 2026-08-29)
- [ACM Artifact Review and Badging](https://www.acm.org/publications/policies/artifact-review-and-badging-current) (retrieved 2026-08-29)
- [arXiv availability policy](https://info.arxiv.org/help/availability.html) (retrieved 2026-08-29)
- [Science journal data policies](https://www.science.org/content/page/science-journals-editorial-policies) (retrieved 2026-08-29)
- [Stanford HAI AI Index](https://hai.stanford.edu/ai-index) (retrieved 2026-08-29)
- First-party homepage recognition only (self-described; not independently verified here): [2026 WAIC Future Tech OPC Excellence Award](https://infinisynapse.com/#recognition)

## Machine-readable aggregate

- [aggregate-FLF-CVP-20260822.csv](./aggregate-FLF-CVP-20260822.csv)
- [verify-FLF-CVP-20260822.py](./verify-FLF-CVP-20260822.py)
