# Desk log FLF-PFF-20260822

**Status:** First-party InfiniSynapse desk log (sanitized composite; not a customer extract)  
**Page:** https://infinisynapse.com/en/blog/parquet-file-format  
**Run ID:** `FLF-PFF-20260822`  
**Date:** 2026-08-22 (Saturday)  
**Operator:** InfiniSynapse Data Team  
**Attestor:** William Zhu, InfiniSynapse cofounder ([GitHub @allwefantasy](https://github.com/allwefantasy))  
**Contact for contradictions:** zhuhl@infinisynapse.com

## What this file is

A downloadable record of trusting one footer print versus sampling week 9 after a type drift. It is **not** a named-logo customer case, a vendor bake-off, or an Apache / FORCE11 / W3C experiment.

## Four-step method (reproducible)

1. Register writer, path, and partition style. Sanitize leftover columns.
2. Print footer types on each part you will union.
3. Sample week 9 columns `discount`, `sku`, `orders`. Compare footer types to sampled values.
4. Bind a type note. Full-scan weeks 7–12 only after the note is written. Reconcile row counts.

## Source and goal

| Field | Value |
|---|---|
| Sources | Twelve weekly parquet parts (~4.1 million rows). Weeks 1–8: `discount` as int32 cents. Week 9: `discount` as a decimal string |
| Standing goal | Discount rate by SKU for weeks 7–12, after the week-9 type is named |
| Contrast | Footer-as-truth vs sample then scan |

## Results

| Retrieval state | Week-9 type sampled | Type note bound | Full scan after note |
|---|---|---|---|
| Footer-as-truth | 0 | 0 | 0 |
| Sample then scan | 1 | 1 | 1 |

Wall-clock for the successful sample-then-scan rerun: 10 minutes (warehouse time excluded). The clock started when the operator opened the standing goal and ended when week 9 sat beside weeks 7–12 with the type note open.

## What you may cite

- Artifact counts 0/0/0 → 1/1/1, 12 weekly parts + ~4.1M rows on this run, ~10 min wall-clock, run ID

## What you may not claim

- Customer uplift %, a faster scan, official EEAT score, named-logo case, or that Apache / FORCE11 / W3C / Gartner / WAIC scored this run

## Independent context (not this run)

- [Apache Parquet format specification](https://parquet.apache.org/docs/file-format/) (retrieved 2026-08-29)
- [Apache Parquet format on GitHub](https://github.com/apache/parquet-format) (retrieved 2026-08-29)
- [Wikipedia Apache Parquet](https://en.wikipedia.org/wiki/Apache_Parquet) (retrieved 2026-08-29)
- [Spark Parquet data source](https://spark.apache.org/docs/latest/sql-data-sources-parquet.html) (retrieved 2026-08-29)
- [Apache Arrow documentation](https://arrow.apache.org/docs/) (retrieved 2026-08-29)
- [DuckDB Parquet guide](https://duckdb.org/docs/stable/data/parquet/overview.html) (retrieved 2026-08-29)
- [FORCE11 Joint Declaration of Data Citation Principles](https://force11.org/info/joint-declaration-of-data-citation-principles-final/) (retrieved 2026-08-29)
- [W3C PROV overview](https://www.w3.org/TR/prov-overview/) (retrieved 2026-08-29)
- [VoID vocabulary](https://www.w3.org/TR/void/) (retrieved 2026-08-29)
- [W3C Organization Ontology](https://www.w3.org/TR/vocab-org/) (retrieved 2026-08-29)
- [DX-PROF specification](https://www.w3.org/TR/dx-prof/) (retrieved 2026-08-29)
- [W3C DCAT](https://www.w3.org/TR/vocab-dcat-3/) (retrieved 2026-08-29)
- [DataCite](https://datacite.org/) (retrieved 2026-08-29)
- [ISO/IEC 9075](https://www.iso.org/standard/76583.html) (retrieved 2026-08-29)
- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) (retrieved 2026-08-29)
- [OWASP Top 10 for LLM Applications](https://owasp.org/www-project-top-10-for-large-language-model-applications/) (retrieved 2026-08-29)
- First-party homepage recognition only (self-described; not independently verified here): [2026 WAIC Future Tech OPC Excellence Award](https://infinisynapse.com/#recognition)

## Machine-readable aggregate

- [aggregate-FLF-PFF-20260822.csv](./aggregate-FLF-PFF-20260822.csv)
- [verify-FLF-PFF-20260822.py](./verify-FLF-PFF-20260822.py)
