# Desk log FLF-SPF-20260822

**Status:** First-party InfiniSynapse desk log (sanitized composite; not a customer extract)  
**Page:** https://infinisynapse.com/en/blog/sample-parquet-file  
**Run ID:** `FLF-SPF-20260822`  
**Date:** 2026-08-22 (Saturday)  
**Operator:** InfiniSynapse Data Team  
**Attestor:** William Zhu, InfiniSynapse cofounder ([GitHub @allwefantasy](https://github.com/allwefantasy))  
**Contact for contradictions:** zhuhl@infinisynapse.com

## What this file is

A downloadable record of shipping a sample rate versus closing the same goal on a full scan. It is **not** a named-logo customer case, a vendor bake-off, or an IFLA / OCLC / ISO experiment.

## Four-step method (reproducible)

1. Sanitize leftover columns. Write one sentence: columns, window, budget, ugly partition.
2. Run that sample on `sku`, `orders`, `returns`.
3. Run the same goal as a full scan. Same grain, same filters, same denominator.
4. Label both rates. Do not average them. Reconcile the gap.

## Source and goal

| Field | Value |
|---|---|
| Sources | Twelve weekly parquet parts (~4.3 million rows) |
| Standing goal | Return rate by SKU, denominator = orders |
| Contrast | Ship the sample (weeks 7–8 only) vs close the scan (weeks 7–12) |

## Results

| Retrieval state | Sample plan written | Same-goal full scan | Rates labeled |
|---|---|---|---|
| Ship the sample | 0 | 0 | 0 |
| Close the scan | 1 | 1 | 1 |

On this desk run the sample rate (weeks 7–8) was 4.1% and the scan rate (weeks 7–12) was 6.8% on the same denominator. The gap was the ugly weeks, not a model.

Wall-clock for the successful close: 10 minutes (warehouse time excluded). The clock started when the operator opened the standing goal and ended when both rates sat side by side with the labels open.

## What you may cite

- Artifact counts 0/0/0 → 1/1/1, 12 weekly parts + ~4.3M rows, 4.1% vs 6.8% on this run, ~10 min wall-clock, run ID

## What you may not claim

- Customer uplift %, a faster sample, official EEAT score, named-logo case, or that IFLA / OCLC / ISO / Gartner / WAIC scored this run

## Independent context (not this run)

- [IFLA Library Reference Model](https://repository.ifla.org/items/214c74cb-c075-4428-a138-39f8d06c55aa) (retrieved 2026-08-29)
- [OCLC Research publications](https://www.oclc.org/research/publications.html) (retrieved 2026-08-29)
- [ISO 26324 (DOI)](https://www.iso.org/standard/53698.html) (retrieved 2026-08-29)
- [ANSI/NISO Z39.29 bibliographic references](https://www.niso.org/publications/ansiniso-z3929-2005-r2010-bibliographic-references) (retrieved 2026-08-29)
- [Crossref metadata best practices](https://www.crossref.org/documentation/principles-practices/best-practices/) (retrieved 2026-08-29)
- [Apache Parquet format specification](https://parquet.apache.org/docs/file-format/) (retrieved 2026-08-29)
- [Apache Parquet format on GitHub](https://github.com/apache/parquet-format) (retrieved 2026-08-29)
- [Wikipedia Sampling (statistics)](https://en.wikipedia.org/wiki/Sampling_(statistics)) (retrieved 2026-08-29)
- [Wikipedia Sampling bias](https://en.wikipedia.org/wiki/Sampling_bias) (retrieved 2026-08-29)
- [Spark DataFrame.sample](https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.DataFrame.sample.html) (retrieved 2026-08-29)
- [Apache Arrow documentation](https://arrow.apache.org/docs/) (retrieved 2026-08-29)
- [DuckDB Parquet guide](https://duckdb.org/docs/stable/data/parquet/overview.html) (retrieved 2026-08-29)
- [W3C DCAT](https://www.w3.org/TR/vocab-dcat-3/) (retrieved 2026-08-29)
- [DataCite](https://datacite.org/) (retrieved 2026-08-29)
- [ISO/IEC 9075](https://www.iso.org/standard/76583.html) (retrieved 2026-08-29)
- First-party homepage recognition only (self-described; not independently verified here): [2026 WAIC Future Tech OPC Excellence Award](https://infinisynapse.com/#recognition)

## Machine-readable aggregate

- [aggregate-FLF-SPF-20260822.csv](./aggregate-FLF-SPF-20260822.csv)
- [verify-FLF-SPF-20260822.py](./verify-FLF-SPF-20260822.py)
