From source rows to a defensible visualization
Reproduce the Kunshan–Venice case, inspect every transformation, and then replace the instructor example with your own domain, data, tasks, visual idioms, algorithms, and evidence.
Instructor example—not a submission topic. Students must replace the domain, community, dataset, tasks, idioms, evidence, and sources with project-specific choices.
How to use this notebook¶
| 1 · Build verified data | 2 · Reproduce the views | 3 · Critique and redesign |
|---|---|---|
| Load nine traceable GeoNames rows, verify their fields, and classify data plus tasks. | Create two interactive maps, a population comparison, a source inspector, and an evidence boundary. | Apply the four-level pipeline, explore all 38 idioms and 72 tools, audit governance, and generate replacement prompts. |
Run cells from top to bottom. The default mode uses the exact verified subset embedded in this notebook. Switch the data-mode control to re-query the pinned Parquet source when you want to audit the extraction itself.
# @title 0 · Setup the Colab environment
import importlib.util
import subprocess
import sys
required = {"duckdb": "duckdb>=1.2,<2", "folium": "folium>=0.18,<2"}
missing = [spec for module, spec in required.items() if importlib.util.find_spec(module) is None]
if missing:
subprocess.check_call([sys.executable, "-m", "pip", "install", "-q", *missing])
import json
import shutil
from pathlib import Path
import duckdb
import folium
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from IPython.display import HTML, Markdown, display
OUTPUT_DIR = Path("infovis_outputs")
OUTPUT_DIR.mkdir(exist_ok=True)
COLORS = {
"ink": "#18313B",
"muted": "#63777D",
"river": "#14798A",
"jiangnan": "#C65F49",
"venice": "#6D6397",
"gold": "#E4B94D",
"line": "#CCD8D5",
"paper": "#F7F5EF",
}
plt.rcParams.update({
"font.family": "DejaVu Sans",
"axes.titleweight": "bold",
"axes.edgecolor": COLORS["line"],
"axes.labelcolor": COLORS["muted"],
"xtick.color": COLORS["muted"],
"ytick.color": COLORS["ink"],
"figure.facecolor": "white",
"axes.facecolor": "white",
})
print("Environment ready · outputs will be written to", OUTPUT_DIR.resolve())
1 · Build the verified data¶
Three moves¶
- Trace the source: Hugging Face conversion → pinned Parquet revision → original GeoNames export.
- Validate the subset: exact identifiers, row count, feature classes, coordinates, and nonzero population fields.
- Abstract before drawing: name the dataset/attribute types and express each task as an action + target.
1.1 Source contract¶
- Dataset: do-me/Geonames on Hugging Face
- Pinned revision:
00c727680a086eaae649043a6ce660a88e57ebea - Original producer/export: GeoNames data dump
- Dataset-card license tag: CC BY 4.0
The nine rows are purposively selected for teaching. They are not a probability sample, and feature counts cannot be interpreted as regional prevalence.
# @title 1.2 · Load instantly or re-query the pinned source
DATA_MODE = "Use embedded verified subset" # @param ["Use embedded verified subset", "Re-query pinned Parquet"]
DATASET_SHA = "00c727680a086eaae649043a6ce660a88e57ebea"
PARQUET_URL = (
"https://huggingface.co/datasets/do-me/Geonames/resolve/"
+ DATASET_SHA + "/geonames_23_03_2025.parquet"
)
FOCUS = {
1785623: {"region": "Jiangnan", "role": "focus city"},
1784074: {"region": "Jiangnan", "role": "water-town settlement"},
1812915: {"region": "Jiangnan", "role": "lake"},
1793666: {"region": "Jiangnan", "role": "lake"},
1886760: {"region": "Jiangnan", "role": "regional city"},
3164603: {"region": "Venice", "role": "focus city"},
3175933: {"region": "Venice", "role": "canal"},
7910672: {"region": "Venice", "role": "canal"},
12172719: {"region": "Venice", "role": "lagoon"},
}
COLUMNS = [
"geonameid", "name", "asciiname", "alternatenames", "latitude",
"longitude", "feature_class", "feature_code", "country_code", "cc2",
"admin1_code", "admin2_code", "admin3_code", "admin4_code",
"population", "elevation", "dem", "timezone", "modification_date",
]
embedded_manifest = json.loads("{\n \"source\": {\n \"dataset_name\": \"do-me/Geonames\",\n \"dataset_card\": \"https://huggingface.co/datasets/do-me/Geonames\",\n \"dataset_sha\": \"00c727680a086eaae649043a6ce660a88e57ebea\",\n \"parquet_url\": \"https://huggingface.co/datasets/do-me/Geonames/resolve/00c727680a086eaae649043a6ce660a88e57ebea/geonames_23_03_2025.parquet\",\n \"upstream\": \"https://download.geonames.org/export/dump/\",\n \"license\": \"CC BY 4.0\",\n \"selection_note\": \"Nine exact GeoNames records selected by stable ID for a paired classroom demonstration. This purposive subset is not exhaustive and feature counts must not be interpreted as regional prevalence.\"\n },\n \"schema\": [\n \"geonameid\",\n \"name\",\n \"asciiname\",\n \"alternatenames\",\n \"latitude\",\n \"longitude\",\n \"feature_class\",\n \"feature_code\",\n \"country_code\",\n \"cc2\",\n \"admin1_code\",\n \"admin2_code\",\n \"admin3_code\",\n \"admin4_code\",\n \"population\",\n \"elevation\",\n \"dem\",\n \"timezone\",\n \"modification_date\"\n ],\n \"records\": [\n {\n \"geonameid\": 1784074,\n \"name\": \"Zhouzhuang\",\n \"asciiname\": \"Zhouzhuang\",\n \"alternatenames\": \"Chou-chuang,Chou-chuang-chen,Chzhouchzhuan,Zhouzhuang,Zhouzhuang Zhen,zhou zhuang,zhou zhuang zhen,\u0427\u0436\u043e\u0443\u0447\u0436\u0443\u0430\u043d,\u5468\u5e84,\u5468\u5e84\u9547\",\n \"latitude\": 31.11788,\n \"longitude\": 120.84427,\n \"feature_class\": \"P\",\n \"feature_code\": \"PPLA4\",\n \"country_code\": \"CN\",\n \"cc2\": null,\n \"admin1_code\": \"04\",\n \"admin2_code\": \"3205\",\n \"admin3_code\": null,\n \"admin4_code\": null,\n \"population\": 22000,\n \"elevation\": null,\n \"dem\": 6,\n \"timezone\": \"Asia/Shanghai\",\n \"modification_date\": \"2021-09-20\",\n \"region\": \"Jiangnan\",\n \"role\": \"water-town settlement\",\n \"source_url\": \"https://www.geonames.org/1784074\"\n },\n {\n \"geonameid\": 1785623,\n \"name\": \"Kunshan\",\n \"asciiname\": \"Kunshan\",\n \"alternatenames\": \"Con Son,C\u00f4n S\u01a1n,K'un-shan-ch'eng,K'un-shan-hsien,KVN,Kan-shan,Kun'shan',Kunsanas,Kunshan,Kunshan Shi,Kun\u0161anas,K\u2019un-shan-ch\u2019eng,K\u2019un-shan-hsien,Yushan,Yushan Zhen,kun shan,kun shan shi,kunsan si,kwnshan,yu shan,yu shan zhen,\u041a\u0443\u043d\u0448\u0430\u043d,\u041a\u0443\u043d\u044c\u0448\u0430\u043d\u044c,\u06a9\u0648\u0646\u0634\u0627\u0646,\u5d11\u5c71,\u5d11\u5c71\u5e02,\u6606\u5c71,\u6606\u5c71\u5e02,\u7389\u5c71,\u7389\u5c71\u9547,\ucfe4\uc0b0 \uc2dc\",\n \"latitude\": 31.37762,\n \"longitude\": 120.95431,\n \"feature_class\": \"P\",\n \"feature_code\": \"PPLA3\",\n \"country_code\": \"CN\",\n \"cc2\": null,\n \"admin1_code\": \"04\",\n \"admin2_code\": \"3205\",\n \"admin3_code\": null,\n \"admin4_code\": null,\n \"population\": 2092496,\n \"elevation\": null,\n \"dem\": 10,\n \"timezone\": \"Asia/Shanghai\",\n \"modification_date\": \"2022-04-01\",\n \"region\": \"Jiangnan\",\n \"role\": \"focus city\",\n \"source_url\": \"https://www.geonames.org/1785623\"\n },\n {\n \"geonameid\": 1793666,\n \"name\": \"Tai Hu\",\n \"asciiname\": \"Tai Hu\",\n \"alternatenames\": \"Danau Taihu,Great Lake,Lac Tai,Lacul Tai,Lago Tai,Lago Taihu,Lake T'ai,Lake Tai,Lake Taihu,Lake T\u2019ai,Llac Taihu,T'ai-wu,Tai Hu,Tai aintzira,Taihu,Taihu Lake,Taijaervi,Taij\u00e4rvi,Taj-to,Taj-t\u00f3,Tajkhu,Tchaj-chu,Thai Ho,Th\u00e1i H\u1ed3,T\u2019ai-wu,Vozera Tajkhu,bhyrt tay,tai ho,tai hu,thale sab thi,\u0412\u043e\u0437\u0435\u0440\u0430 \u0422\u0430\u0439\u0445\u0443,\u0422\u0430\u0439\u0445\u0443,\u0628\u062d\u064a\u0631\u0629 \u062a\u0627\u064a,\u0e17\u0e30\u0e40\u0e25\u0e2a\u0e32\u0e1a\u0e44\u0e17\u0e48,\u0f50\u0f60\u0f7a\u0f0b\u0f67\u0f74\u0f60\u0f74\u0f0b\u0f58\u0f5a\u0f7a\u0f60\u0f74\u0f0d,\u592a\u6e56,\ud0c0\uc774 \ud638\",\n \"latitude\": 31.21649,\n \"longitude\": 120.19814,\n \"feature_class\": \"H\",\n \"feature_code\": \"LK\",\n \"country_code\": \"CN\",\n \"cc2\": null,\n \"admin1_code\": \"04\",\n \"admin2_code\": null,\n \"admin3_code\": null,\n \"admin4_code\": null,\n \"population\": 0,\n \"elevation\": null,\n \"dem\": 0,\n \"timezone\": \"Asia/Shanghai\",\n \"modification_date\": \"2020-12-11\",\n \"region\": \"Jiangnan\",\n \"role\": \"lake\",\n \"source_url\": \"https://www.geonames.org/1793666\"\n },\n {\n \"geonameid\": 1812915,\n \"name\": \"Dianshan Hu\",\n \"asciiname\": \"Dianshan Hu\",\n \"alternatenames\": \"Dianshan Hu,Dianshan Lake,Di\u00e0nsh\u0101n H\u00fa,Tien-shan Hu,dian shan hu,\u6dc0\u5c71\u6e56\",\n \"latitude\": 31.11417,\n \"longitude\": 120.96472,\n \"feature_class\": \"H\",\n \"feature_code\": \"LK\",\n \"country_code\": \"CN\",\n \"cc2\": null,\n \"admin1_code\": \"00\",\n \"admin2_code\": null,\n \"admin3_code\": null,\n \"admin4_code\": null,\n \"population\": 0,\n \"elevation\": null,\n \"dem\": 1,\n \"timezone\": \"Asia/Shanghai\",\n \"modification_date\": \"2024-07-13\",\n \"region\": \"Jiangnan\",\n \"role\": \"lake\",\n \"source_url\": \"https://www.geonames.org/1812915\"\n },\n {\n \"geonameid\": 1886760,\n \"name\": \"Suzhou\",\n \"asciiname\": \"Suzhou\",\n \"alternatenames\": \"SZV,So-chiu-chhi,Soochow,Soutsoou,So\u0358-chiu-chh\u012b,Su-chou,Su-chu-su,Su-ciu,Su-cou,Su-\u010dou,Suchjou,Suchzhou,Sudzhou,Sudzou,Sud\u017eou,Sugouo,Sutsjou,Suzhou,Suzhou Shi,Suzhou i Jiangsu,Su\u011do\u016do,Szucsou,S\u00fb-ch\u00fb-s\u1e73,S\u016d-ci\u016d,To Chau,T\u00f4 Ch\u00e2u,Wu-hsien,cuco,ssujeou si,su cow,su zhou,su zhou shi,sujho'u,suzu,swgw\u02bcw,swjw,swzhw,swzhww,\u03a3\u03bf\u03c5\u03c4\u03c3\u03cc\u03bf\u03c5,\u0421\u0443\u0434\u0436\u043e\u0443,\u0421\u0443\u0447\u0436\u043e\u0443,\u0421\u0443\u0447\u0436\u043e\u045e,\u0421\u0443\u045f\u043e\u0443,\u0421\u04af\u0436\u043e\u0443,\u054d\u0578\u0582\u0579\u056a\u0578\u0578\u0582,\u05e1\u05d5\u05d2\u05d5\u05d0\u05d5,\u0633\u0648\u062c\u0648,\u0633\u0648\u0698\u0648,\u0633\u0648\u0698\u0648\u0648,\u0633\u06c7\u062c\u06c7 \u0634\u06d5\u06be\u0649\u0631\u0649,\u0938\u0942\u091d\u094b\u090a,\u0a38\u0a42\u0a1c\u0a3c\u0a42,\u0b9a\u0bc1\u0b9a\u0bcb,\u0e0b\u0e39\u0e42\u0e08\u0e27,\u82cf\u5dde,\u82cf\u5dde\u5e02,\u8607\u5dde,\u8607\u5dde\u5e02,\uc464\uc800\uc6b0 \uc2dc\",\n \"latitude\": 31.30408,\n \"longitude\": 120.59538,\n \"feature_class\": \"P\",\n \"feature_code\": \"PPLA2\",\n \"country_code\": \"CN\",\n \"cc2\": null,\n \"admin1_code\": \"04\",\n \"admin2_code\": \"3205\",\n \"admin3_code\": null,\n \"admin4_code\": null,\n \"population\": 6715559,\n \"elevation\": null,\n \"dem\": 10,\n \"timezone\": \"Asia/Shanghai\",\n \"modification_date\": \"2024-03-26\",\n \"region\": \"Jiangnan\",\n \"role\": \"regional city\",\n \"source_url\": \"https://www.geonames.org/1886760\"\n },\n {\n \"geonameid\": 3164603,\n \"name\": \"Venice\",\n \"asciiname\": \"Venice\",\n \"alternatenames\": \"Benatky,Benetia,Benetke,Benezia,Ben\u00e1tky,Feneyjar,Mleci,V'nise,VCE,Velence,Venecia,Venecia - Venezia,Venecija,Venecio,Venedeg,Venedig,Venedik,Venediku,Venesia,Venesiya,Venetia,Venetie,Venetik,Veneti\u00eb,Venetsia,Veneza,Venezia,Venezsia,Vene\u021bia,Venice,Venies,Venise,Venizia,Ven\u00e8cia,Ven\u00e8sia,Vignesie,Vinezzia,Wenecja,albndqyt,an Veineis,an Vein\u00e9is,benechia,beniseu,benisu,venetsia,vuenetsuia,vuenisu,wei ni si,wnyz,wnzyh,\u0392\u03b5\u03bd\u03b5\u03c4\u03af\u03b1,\u0412\u0435\u043d\u0435\u0446\u0438\u044f,\u0412\u0435\u043d\u0435\u0446\u0438\u0458\u0430,\u0412\u0435\u043d\u0435\u0446\u0456\u044f,\u054e\u0565\u0576\u0565\u057f\u056b\u056f,\u05d5\u05e0\u05e6\u05d9\u05d4,\u0627\u0644\u0628\u0646\u062f\u0642\u064a\u0629,\u0648\u0646\u06cc\u0632,\u06cb\u06d0\u0646\u0649\u062a\u0633\u0649\u064a\u06d5,\u10d5\u10d4\u10dc\u10d4\u10ea\u10d8\u10d0,\u30d9\u30cb\u30b9,\u30f4\u30a7\u30cb\u30b9,\u30f4\u30a7\u30cd\u30c4\u30a3\u30a2,\u5a01\u5c3c\u65af,\ubca0\ub124\uce58\uc544,\ubca0\ub2c8\uc2a4\",\n \"latitude\": 45.43713,\n \"longitude\": 12.33265,\n \"feature_class\": \"P\",\n \"feature_code\": \"PPLA\",\n \"country_code\": \"IT\",\n \"cc2\": null,\n \"admin1_code\": \"20\",\n \"admin2_code\": \"VE\",\n \"admin3_code\": \"027042\",\n \"admin4_code\": null,\n \"population\": 51298,\n \"elevation\": 2.0,\n \"dem\": 5,\n \"timezone\": \"Europe/Rome\",\n \"modification_date\": \"2025-02-08\",\n \"region\": \"Venice\",\n \"role\": \"focus city\",\n \"source_url\": \"https://www.geonames.org/3164603\"\n },\n {\n \"geonameid\": 3175933,\n \"name\": \"Canal Grande\",\n \"asciiname\": \"Canal Grande\",\n \"alternatenames\": \"Bueyuek Kanal,B\u00fcy\u00fck Kanal,Canal Grande,Canal Grando,Golem kanal,Gran Canal,Gran Canal de Venecia,Grand Canal,Grand-kanal,Granda Kanalo de Venecio,Grande Canal de Veneza,Kanal Grande,Kanal Qrande,Lielais kanals,Lielais kan\u0101ls,alqnal alkbyr,da yun he,kanal geulande,\u0413\u043e\u043b\u0435\u043c \u043a\u0430\u043d\u0430\u043b,\u0413\u0440\u0430\u043d\u0434-\u043a\u0430\u043d\u0430\u043b,\u041a\u0430\u043d\u0430\u043b \u0413\u0440\u0430\u043d\u0434\u0435,\u05d4\u05ea\u05e2\u05dc\u05d4 \u05d4\u05d2\u05d3\u05d5\u05dc\u05d4,\u0627\u0644\u0642\u0646\u0627\u0644 \u0627\u0644\u0643\u0628\u064a\u0631,\u30ab\u30ca\u30eb\u30fb\u30b0\u30e9\u30f3\u30c7,\u5927\u8fd0\u6cb3,\uce74\ub0a0 \uadf8\ub780\ub370\",\n \"latitude\": 45.43598,\n \"longitude\": 12.33052,\n \"feature_class\": \"H\",\n \"feature_code\": \"CNL\",\n \"country_code\": \"IT\",\n \"cc2\": null,\n \"admin1_code\": \"20\",\n \"admin2_code\": \"VE\",\n \"admin3_code\": \"027042\",\n \"admin4_code\": null,\n \"population\": 0,\n \"elevation\": null,\n \"dem\": 6,\n \"timezone\": \"Europe/Rome\",\n \"modification_date\": \"2018-04-05\",\n \"region\": \"Venice\",\n \"role\": \"canal\",\n \"source_url\": \"https://www.geonames.org/3175933\"\n },\n {\n \"geonameid\": 7910672,\n \"name\": \"Canal Grande di Murano\",\n \"asciiname\": \"Canal Grande di Murano\",\n \"alternatenames\": null,\n \"latitude\": 45.45602,\n \"longitude\": 12.35462,\n \"feature_class\": \"H\",\n \"feature_code\": \"CNL\",\n \"country_code\": \"IT\",\n \"cc2\": null,\n \"admin1_code\": \"20\",\n \"admin2_code\": \"VE\",\n \"admin3_code\": \"027042\",\n \"admin4_code\": null,\n \"population\": 0,\n \"elevation\": null,\n \"dem\": 2,\n \"timezone\": \"Europe/Rome\",\n \"modification_date\": \"2011-07-29\",\n \"region\": \"Venice\",\n \"role\": \"canal\",\n \"source_url\": \"https://www.geonames.org/7910672\"\n },\n {\n \"geonameid\": 12172719,\n \"name\": \"Venetian Lagoon\",\n \"asciiname\": \"Venetian Lagoon\",\n \"alternatenames\": \"Venetian Lagoon\",\n \"latitude\": 45.3662,\n \"longitude\": 12.25136,\n \"feature_class\": \"H\",\n \"feature_code\": \"LGN\",\n \"country_code\": \"IT\",\n \"cc2\": null,\n \"admin1_code\": \"20\",\n \"admin2_code\": \"VE\",\n \"admin3_code\": \"027023\",\n \"admin4_code\": null,\n \"population\": 0,\n \"elevation\": null,\n \"dem\": -2,\n \"timezone\": \"Europe/Rome\",\n \"modification_date\": \"2020-10-26\",\n \"region\": \"Venice\",\n \"role\": \"lagoon\",\n \"source_url\": \"https://www.geonames.org/12172719\"\n }\n ]\n}")
if DATA_MODE == "Use embedded verified subset":
manifest = embedded_manifest
records = pd.DataFrame(manifest["records"])
extraction_note = "Loaded the embedded copy of the verified nine-row manifest."
else:
aliases = ", ".join(f'"{i}" AS {name}' for i, name in enumerate(COLUMNS))
identifiers = ", ".join(map(str, FOCUS))
query = f"""
SELECT {aliases}
FROM read_parquet('{PARQUET_URL}')
WHERE "0" IN ({identifiers})
ORDER BY "0"
"""
connection = duckdb.connect()
records = connection.execute(query).df()
for row_index, geonameid in records["geonameid"].astype(int).items():
records.loc[row_index, ["region", "role"]] = list(FOCUS[geonameid].values())
records["source_url"] = records["geonameid"].map(
lambda value: f"https://www.geonames.org/{int(value)}"
)
manifest = embedded_manifest
extraction_note = "Re-queried nine IDs from the pinned Hugging Face Parquet source."
records["geonameid"] = records["geonameid"].astype(int)
records["population"] = pd.to_numeric(records["population"], errors="coerce").fillna(0).astype(int)
records["latitude"] = pd.to_numeric(records["latitude"], errors="raise")
records["longitude"] = pd.to_numeric(records["longitude"], errors="raise")
records.to_csv(OUTPUT_DIR / "geonames_water_towns.csv", index=False)
print(extraction_note)
records[["name", "region", "role", "feature_class", "feature_code", "population"]]
# @title 1.3 · Run the integrity contract
expected_ids = set(FOCUS)
observed_ids = set(records["geonameid"])
populated = records.query("feature_class == 'P' and population > 0")
checks = {
"Exactly nine rows": len(records) == 9,
"All stable IDs found": observed_ids == expected_ids,
"No duplicate IDs": records["geonameid"].is_unique,
"Only H/P feature classes": set(records["feature_class"]) == {"H", "P"},
"Coordinates are valid": (
records["latitude"].between(-90, 90).all()
and records["longitude"].between(-180, 180).all()
),
"Four nonzero settlement values": len(populated) == 4,
}
assert all(checks.values()), checks
display(pd.DataFrame({"check": checks.keys(), "passed": checks.values()}))
print("Integrity contract passed.")
1.4 Course cheat sheets¶
These are the same course figures used in the tutorial. Click the source links for the editable/high-resolution originals.
| What? Data and attributes | Why? Actions and targets |
|---|---|
| Figure 2.1 source PDF | Figure 3.1 source PDF |
Source: Tamara Munzner, Visualization Analysis and Design (2014), illustrations by Eamonn Maguire, CC BY 4.0.
# @title 1.5 · Name the data, attributes, actions, and targets
attribute_contract = pd.DataFrame([
["Table", "Each row is one selected GeoNames record", "Dataset type"],
["Geometry", "Longitude + latitude become point positions only after spatial encoding", "Derived dataset type"],
["Identifier", "geonameid", "Attribute role"],
["Categorical", "name, feature_class/code, country, timezone, region, role", "Attribute type"],
["Quantitative", "latitude, longitude, population, elevation, DEM", "Attribute type"],
], columns=["class", "case abstraction", "course vocabulary"])
task_contract = pd.DataFrame([
["Locate", "selected named features", "Search → spatial target"],
["Browse", "available source fields", "Search → item/attribute target"],
["Compare", "population for four settlement rows", "Query → quantitative attribute target"],
["Identify", "unsupported cultural/environmental claims", "Query → evidence boundary"],
], columns=["action", "target", "action + target expression"])
display(attribute_contract.style.hide(axis="index"))
display(task_contract.style.hide(axis="index"))
# @title 1.6 · Inspect the exact source rows
source_columns = [
"geonameid", "name", "region", "role", "feature_class", "feature_code",
"latitude", "longitude", "population", "modification_date", "source_url",
]
source_view = records[source_columns].sort_values(["region", "feature_class", "name"])
display(
source_view.style
.format({"latitude": "{:.5f}", "longitude": "{:.5f}", "population": "{:,.0f}"})
.hide(axis="index")
)
2 · Reproduce the demo views¶
Three linked views¶
- Map: two local geographic views answer where without forcing labels to overlap.
- Population: aligned log position and exact labels answer how much more accurately than area alone.
- Records + limits: the source inspector and evidence table prevent geographic context from becoming a cultural conclusion.
# @title 2.1 · Build the two interactive maps
MAP_FILTER = "All records" # @param ["All records", "Water features", "Settlements"]
filter_code = {"All records": None, "Water features": "H", "Settlements": "P"}[MAP_FILTER]
map_rows = records if filter_code is None else records.query("feature_class == @filter_code")
settlement_values = records.query("feature_class == 'P' and population > 0")["population"]
log_low, log_high = np.log10(settlement_values.min()), np.log10(settlement_values.max())
def graduated_radius(row):
if row.feature_class != "P" or row.population <= 0:
return 7
position = (np.log10(row.population) - log_low) / (log_high - log_low)
return float(8 + 18 * np.sqrt(max(0, position)))
def region_map(region):
all_region = records.query("region == @region")
shown = map_rows.query("region == @region")
center = [all_region["latitude"].mean(), all_region["longitude"].mean()]
view = folium.Map(
location=center, tiles="OpenStreetMap", control_scale=True,
zoom_control=True, scrollWheelZoom=False, width="100%", height=420,
)
for row in shown.itertuples():
water = row.feature_class == "H"
color = COLORS["river"] if water else (
COLORS["jiangnan"] if row.region == "Jiangnan" else COLORS["venice"]
)
population_text = f"{row.population:,}" if row.population > 0 else "not applicable / 0"
popup = folium.Popup(
f"<b>{row.name}</b><br>Class/code: {row.feature_class}/{row.feature_code}"
f"<br>Population field: {population_text}"
f'<br><a href="{row.source_url}" target="_blank">Open GeoNames record</a>',
max_width=280,
)
folium.CircleMarker(
[row.latitude, row.longitude],
radius=graduated_radius(row),
color=COLORS["river"] if water else "white",
weight=3 if water else 2,
fill=True,
fill_color=COLORS["river"] if water else color,
fill_opacity=0.35 if water else 0.82,
tooltip=f"{row.name} · {row.feature_class}/{row.feature_code}",
popup=popup,
).add_to(view)
view.fit_bounds([
[all_region["latitude"].min(), all_region["longitude"].min()],
[all_region["latitude"].max(), all_region["longitude"].max()],
], padding=(24, 24))
return view
jiangnan_map = region_map("Jiangnan")
venice_map = region_map("Venice")
jiangnan_map.save(OUTPUT_DIR / "jiangnan_map.html")
venice_map.save(OUTPUT_DIR / "venice_map.html")
display(HTML(
'<div style="display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:14px;">'
'<div><h3 style="font-family:Georgia;color:#18313b;">Kunshan, Suzhou & nearby waters</h3>'
+ jiangnan_map._repr_html_() + '</div>'
'<div><h3 style="font-family:Georgia;color:#18313b;">Venice & its lagoon</h3>'
+ venice_map._repr_html_() + '</div></div>'
'<p style="color:#63777d;font-size:13px;">'
'Map bubbles are a graduated cue; use the comparison chart below for exact values. '
'Basemap © OpenStreetMap contributors.</p>'
))
# @title 2.2 · Draw and export the population comparison
population = (
records.query("feature_class == 'P' and population > 0")
.sort_values("population", ascending=True)
.copy()
)
population["color"] = population["region"].map({
"Jiangnan": COLORS["jiangnan"], "Venice": COLORS["venice"]
})
population["bubble_area"] = 90 + 560 * np.sqrt(
population["population"] / population["population"].max()
)
fig, ax = plt.subplots(figsize=(10.5, 5.7))
y = np.arange(len(population))
ax.hlines(y, 1e4, population["population"], color=COLORS["line"], linewidth=2, zorder=1)
ax.scatter(
population["population"], y,
s=population["bubble_area"], c=population["color"],
edgecolor="white", linewidth=2.5, zorder=3,
)
for index, row in enumerate(population.itertuples()):
ax.annotate(
f"{row.population:,}",
(row.population, index),
xytext=(-13 if row.population > 1_000_000 else 13, 0),
textcoords="offset points",
ha="right" if row.population > 1_000_000 else "left",
va="center", fontsize=10, fontweight="bold", color=COLORS["ink"],
)
labels = [f"{row.name}\n{row.region} · {row.role}" for row in population.itertuples()]
ax.set_yticks(y, labels)
ax.set_xscale("log")
ax.set_xlim(1e4, 1.15e7)
ax.set_xticks([1e4, 1e5, 1e6, 1e7], ["10K", "100K", "1M", "10M"])
ax.grid(axis="x", color=COLORS["line"], linewidth=.8)
ax.set_axisbelow(True)
ax.set_xlabel("Population field recorded by GeoNames · logarithmic scale")
fig.suptitle(
"Four settlement records on one comparable scale",
x=0.235, y=0.965, ha="left", fontsize=19, fontweight="bold", color=COLORS["ink"],
)
fig.text(
0.235, 0.91,
"Position supports comparison; bubble area is secondary; exact values are printed.",
ha="left", color=COLORS["muted"], fontsize=10,
)
for spine in ["top", "right", "left"]:
ax.spines[spine].set_visible(False)
ax.tick_params(axis="y", length=0)
fig.subplots_adjust(left=0.235, right=0.97, bottom=0.15, top=0.84)
for extension in ["png", "svg", "pdf"]:
fig.savefig(
OUTPUT_DIR / f"population_comparison.{extension}",
dpi=220 if extension == "png" else None,
bbox_inches="tight",
)
plt.show()
display(Markdown(
"**Comparability note.** These are the values in four selected GeoNames rows. "
"The source does not provide harmonized reference dates or definitions, so this "
"figure is descriptive of the recorded field—not a synchronized census comparison."
))
# @title 2.3 · Inspect one record without leaving the notebook
SELECTED_RECORD = "Zhouzhuang" # @param ["Zhouzhuang", "Kunshan", "Suzhou", "Tai Hu", "Dianshan Hu", "Venice", "Canal Grande", "Canal Grande di Murano", "Venetian Lagoon"]
selected = records.loc[records["name"].eq(SELECTED_RECORD)].iloc[0]
detail = pd.DataFrame({
"field": [
"GeoNames ID", "role", "feature class / code", "coordinates",
"population field", "elevation / DEM", "timezone", "modified", "source",
],
"value": [
selected.geonameid, selected.role,
f"{selected.feature_class} / {selected.feature_code}",
f"{selected.latitude:.5f}, {selected.longitude:.5f}",
f"{selected.population:,}" if selected.population else "not applicable / 0",
f"{selected.elevation} / {selected.dem}", selected.timezone,
selected.modification_date, selected.source_url,
],
})
display(detail.style.hide(axis="index"))
# @title 2.4 · State the evidence boundary
evidence_boundary = pd.DataFrame([
[
"Supported",
"Named geographic context",
"Names, coordinates, feature class/code, and selected administrative fields",
"Locate these nine rows and inspect their available attributes",
],
[
"Partly supported",
"Population and water-system framing",
"Population and class-H fields exist; reference dates, definitions, connectivity, condition, use, and change do not",
"Describe recorded fields with explicit limitations",
],
[
"Not supported",
"Cultural meaning or community need",
"No evidence of heritage value, exchange, authority, tourism pressure, or a historical route",
"Do not infer; add authoritative and participatory evidence",
],
], columns=["status", "claim area", "what the rows contain", "permitted use"])
display(evidence_boundary.style.hide(axis="index"))
# @title 3.1 · Apply the four-level pipeline to the case
pipeline = pd.DataFrame([
[
"Domain",
"Whose real question should the artifact serve?",
"Intercultural inquiry around two water systems; local need remains an open question",
"Keep the interface framed as a starting point, not a community contribution",
],
[
"Data / task",
"Which dataset/attribute types and action + target pairs are present?",
"Table → point geometry; categorical + quantitative attributes; locate, browse, compare, identify",
"Limit comparison to recorded fields and mark absent evidence",
],
[
"Idiom",
"Which form follows from the task?",
"Two maps for locate; log bubble-lollipop for compare; table for inspect",
"Do not rely on map-circle area for exact comparison",
],
[
"Algorithm",
"Which transformations and tools produce the idiom?",
"Pinned-ID filter → field mapping → validation → Folium maps + Matplotlib chart",
"Export source mapping and PNG/SVG/PDF; retain failure and comparability notes",
],
], columns=["level", "question", "worked-case answer", "visible design consequence"])
display(pipeline.style.hide(axis="index"))
3.2 Idiom decision tree¶
Choose an idiom only after naming the action and target. The next cell includes every option shown in the tutorial:
- Compare: values/ranks, change/groups, composition/profiles
- Understand variation: distributions, relationships/density, uncertainty/models
- Reveal structure: hierarchy/flow, cycles/opposition, arranged tables
For this case: Locate → map and Compare → log-scaled bubble-lollipop. Treemap is rejected because these rows contain no hierarchy; population pyramid is rejected because they contain no opposing populations over shared ordered categories.
# @title 3.3 · Explore all 38 demonstrated idioms
IDIOM_ROOT = "Compare" # @param ["Compare", "Understand variation", "Reveal structure"]
idiom_catalog = pd.DataFrame([{'root': 'Compare', 'family': 'Values and ranks', 'idiom': 'Sorted bar', 'selection_cue': 'Compare magnitudes on a shared zero baseline.', 'caution': ''}, {'root': 'Compare', 'family': 'Values and ranks', 'idiom': 'Dot plot', 'selection_cue': 'Compare positions with less ink than bars.', 'caution': ''}, {'root': 'Compare', 'family': 'Values and ranks', 'idiom': 'Lollipop', 'selection_cue': 'Emphasize endpoints while retaining a baseline.', 'caution': ''}, {'root': 'Compare', 'family': 'Change and groups', 'idiom': 'Slope graph', 'selection_cue': 'Compare two time points across groups.', 'caution': ''}, {'root': 'Compare', 'family': 'Change and groups', 'idiom': 'Dumbbell', 'selection_cue': 'Compare paired values and the gap between them.', 'caution': ''}, {'root': 'Compare', 'family': 'Change and groups', 'idiom': 'Line chart', 'selection_cue': 'Follow ordered change; do not connect unordered categories.', 'caution': ''}, {'root': 'Compare', 'family': 'Change and groups', 'idiom': 'Small multiples', 'selection_cue': 'Compare repeated panels with aligned scales.', 'caution': ''}, {'root': 'Compare', 'family': 'Composition and profiles', 'idiom': 'Stacked bar', 'selection_cue': 'Compare totals and a limited number of parts.', 'caution': ''}, {'root': 'Compare', 'family': 'Composition and profiles', 'idiom': '100% stacked', 'selection_cue': 'Compare proportions when totals are secondary.', 'caution': ''}, {'root': 'Compare', 'family': 'Composition and profiles', 'idiom': 'Donut', 'selection_cue': 'Use only for a few parts; angles are imprecise.', 'caution': 'Use sparingly'}, {'root': 'Compare', 'family': 'Composition and profiles', 'idiom': 'Aligned profile', 'selection_cue': 'Compare many measures on shared aligned axes.', 'caution': ''}, {'root': 'Compare', 'family': 'Composition and profiles', 'idiom': 'Radar', 'selection_cue': 'Only for a small, common scale; aligned profiles are often clearer.', 'caution': 'Requires justification'}, {'root': 'Understand variation', 'family': 'One variable', 'idiom': 'Histogram', 'selection_cue': 'Show binned distribution shape; test bin sensitivity.', 'caution': ''}, {'root': 'Understand variation', 'family': 'One variable', 'idiom': 'KDE', 'selection_cue': 'Show smoothed density; disclose bandwidth.', 'caution': ''}, {'root': 'Understand variation', 'family': 'One variable', 'idiom': 'ECDF', 'selection_cue': 'Show every observation without bin or bandwidth choice.', 'caution': ''}, {'root': 'Understand variation', 'family': 'One variable', 'idiom': 'Box plot', 'selection_cue': 'Compare robust summaries; show raw data when sample size matters.', 'caution': ''}, {'root': 'Understand variation', 'family': 'One variable', 'idiom': 'Violin', 'selection_cue': 'Compare distribution shape; include scale and sample context.', 'caution': ''}, {'root': 'Understand variation', 'family': 'One variable', 'idiom': 'Raincloud', 'selection_cue': 'Combine density, summary, and observations.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Relationships and density', 'idiom': 'Scatterplot', 'selection_cue': 'Inspect two quantitative attributes.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Relationships and density', 'idiom': 'Regression view', 'selection_cue': 'Show fitted relationship with uncertainty and assumptions.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Relationships and density', 'idiom': 'Hexbin', 'selection_cue': 'Aggregate dense points to reveal concentration.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Relationships and density', 'idiom': 'Pair plot', 'selection_cue': 'Scan many pairwise relationships; avoid tiny unreadable panels.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Relationships and density', 'idiom': 'Correlation heatmap', 'selection_cue': 'Reveal a matrix pattern after justified ordering.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Relationships and density', 'idiom': '2D contour', 'selection_cue': 'Show a continuous surface without unnecessary 3D perspective.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Uncertainty and models', 'idiom': 'Interval / forest', 'selection_cue': 'Compare estimates and uncertainty on aligned axes.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Uncertainty and models', 'idiom': 'Fan chart', 'selection_cue': 'Show forecast uncertainty expanding through time.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Uncertainty and models', 'idiom': 'Calibration', 'selection_cue': 'Compare predicted probability with observed frequency.', 'caution': ''}, {'root': 'Understand variation', 'family': 'Uncertainty and models', 'idiom': 'Residual view', 'selection_cue': 'Diagnose model error, structure, and outliers.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Hierarchy and flow', 'idiom': 'Treemap', 'selection_cue': 'Encode a genuine part-to-whole hierarchy through nested area.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Hierarchy and flow', 'idiom': 'Sunburst', 'selection_cue': 'Show hierarchical depth radially; labels can become difficult.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Hierarchy and flow', 'idiom': 'Sankey', 'selection_cue': 'Show aggregated flows with conserved quantities.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Hierarchy and flow', 'idiom': 'Alluvial', 'selection_cue': 'Compare category flows across ordered stages.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Cycles and opposition', 'idiom': 'Polar / Burtin', 'selection_cue': 'Use only for genuinely cyclic variables.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Cycles and opposition', 'idiom': 'Population pyramid', 'selection_cue': 'Compare two opposing populations across the same ordered categories.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Arranged tables', 'idiom': 'Ordered heatmap', 'selection_cue': 'Reveal table structure through meaningful row and column order.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Arranged tables', 'idiom': 'Faceted table', 'selection_cue': 'Keep exact values while arranging meaningful groups.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Arranged tables', 'idiom': 'Missingness matrix', 'selection_cue': 'Expose which variables or groups lack evidence.', 'caution': ''}, {'root': 'Reveal structure', 'family': 'Arranged tables', 'idiom': 'UpSet matrix', 'selection_cue': 'Compare intersections when a Venn diagram no longer scales.', 'caution': ''}])
selected_idioms = idiom_catalog.query("root == @IDIOM_ROOT")
display(selected_idioms.style.hide(axis="index"))
print(
f"Showing {len(selected_idioms)} of {len(idiom_catalog)} idioms · "
f"roots: {', '.join(idiom_catalog['root'].unique())}"
)
# @title 3.4 · Search all 72 Python and R visualization tools
LANGUAGE = "Python" # @param ["Python", "R"]
TOOL_CATEGORY = "All" # @param ["All", "grammar", "statistics", "specialized"]
TOOL_SEARCH = "map" # @param {type:"string"}
tool_catalog = pd.DataFrame([{'name': 'Matplotlib', 'language': 'Python', 'category': 'grammar', 'use': 'Foundational static, animated, and publication figures.', 'tags': 'bar line scatter animation', 'official_url': 'https://matplotlib.org/stable/gallery/index.html'}, {'name': 'Seaborn', 'language': 'Python', 'category': 'grammar', 'use': 'Statistical graphics with concise data-aware defaults.', 'tags': 'distribution relational categorical', 'official_url': 'https://seaborn.pydata.org/examples/index.html'}, {'name': 'Plotly', 'language': 'Python', 'category': 'grammar', 'use': 'Interactive browser charts, maps, 3D, and dashboards.', 'tags': 'interactive hover map', 'official_url': 'https://plotly.com/python/'}, {'name': 'Altair', 'language': 'Python', 'category': 'grammar', 'use': 'Declarative Vega-Lite grammar for composable charts.', 'tags': 'declarative grammar interactive', 'official_url': 'https://altair-viz.github.io/gallery/index.html'}, {'name': 'Bokeh', 'language': 'Python', 'category': 'grammar', 'use': 'Interactive linked plots and browser applications.', 'tags': 'interactive linked brushing server', 'official_url': 'https://docs.bokeh.org/en/latest/docs/gallery.html'}, {'name': 'HoloViews', 'language': 'Python', 'category': 'grammar', 'use': 'High-level declarative views across plotting backends.', 'tags': 'declarative linked data', 'official_url': 'https://holoviews.org/reference/index.html'}, {'name': 'hvPlot', 'language': 'Python', 'category': 'grammar', 'use': 'Interactive plotting API for pandas, xarray, and more.', 'tags': 'pandas xarray interactive', 'official_url': 'https://hvplot.holoviz.org/reference/index.html'}, {'name': 'plotnine', 'language': 'Python', 'category': 'grammar', 'use': 'Grammar-of-graphics implementation inspired by ggplot2.', 'tags': 'grammar layers facets', 'official_url': 'https://plotnine.org/gallery.html'}, {'name': 'Lets-Plot', 'language': 'Python', 'category': 'grammar', 'use': 'Grammar-of-graphics plots for notebooks and web output.', 'tags': 'grammar interactive notebooks', 'official_url': 'https://lets-plot.org/python/pages/gallery.html'}, {'name': 'Pygal', 'language': 'Python', 'category': 'grammar', 'use': 'Lightweight SVG charts with browser-friendly output.', 'tags': 'svg browser simple', 'official_url': 'https://www.pygal.org/en/stable/documentation/types/index.html'}, {'name': 'pyecharts', 'language': 'Python', 'category': 'grammar', 'use': 'Python bindings for the Apache ECharts ecosystem.', 'tags': 'echarts interactive web', 'official_url': 'https://gallery.pyecharts.org/'}, {'name': 'Panel', 'language': 'Python', 'category': 'grammar', 'use': 'Compose plots, widgets, and data apps across libraries.', 'tags': 'dashboard widgets app', 'official_url': 'https://panel.holoviz.org/gallery/index.html'}, {'name': 'pandas plotting', 'language': 'Python', 'category': 'statistics', 'use': 'Quick plots directly from Series and DataFrames.', 'tags': 'dataframe quick exploratory', 'official_url': 'https://pandas.pydata.org/docs/user_guide/visualization.html'}, {'name': 'statsmodels graphics', 'language': 'Python', 'category': 'statistics', 'use': 'Regression, diagnostic, time-series, and model plots.', 'tags': 'regression residual diagnostic', 'official_url': 'https://www.statsmodels.org/stable/graphics.html'}, {'name': 'scikit-learn Displays', 'language': 'Python', 'category': 'statistics', 'use': 'Model evaluation displays with estimator integration.', 'tags': 'machine learning calibration roc confusion', 'official_url': 'https://scikit-learn.org/stable/visualizations.html'}, {'name': 'Yellowbrick', 'language': 'Python', 'category': 'statistics', 'use': 'Visual diagnostics for machine-learning workflows.', 'tags': 'machine learning diagnostic residual', 'official_url': 'https://www.scikit-yb.org/en/latest/gallery.html'}, {'name': 'ArviZ', 'language': 'Python', 'category': 'statistics', 'use': 'Exploratory analysis and diagnostics for Bayesian models.', 'tags': 'bayesian posterior interval', 'official_url': 'https://python.arviz.org/en/stable/examples/index.html'}, {'name': 'corner.py', 'language': 'Python', 'category': 'statistics', 'use': 'Multidimensional posterior and parameter distributions.', 'tags': 'bayesian pairplot distribution', 'official_url': 'https://corner.readthedocs.io/en/latest/pages/quickstart/'}, {'name': 'missingno', 'language': 'Python', 'category': 'statistics', 'use': 'Missing-data matrices, bars, heatmaps, and dendrograms.', 'tags': 'missingness data quality', 'official_url': 'https://github.com/ResidentMario/missingno'}, {'name': 'UpSetPlot', 'language': 'Python', 'category': 'statistics', 'use': 'Scalable set-intersection visualization.', 'tags': 'sets intersections upset', 'official_url': 'https://upsetplot.readthedocs.io/en/stable/auto_examples/index.html'}, {'name': 'JoyPy', 'language': 'Python', 'category': 'statistics', 'use': 'Ridgeline distribution plots built on Matplotlib.', 'tags': 'ridgeline density distribution', 'official_url': 'https://github.com/leotac/joypy'}, {'name': 'PtitPrince', 'language': 'Python', 'category': 'statistics', 'use': 'Raincloud plots combining density, box, and points.', 'tags': 'raincloud distribution', 'official_url': 'https://github.com/pog87/PtitPrince'}, {'name': 'SciencePlots', 'language': 'Python', 'category': 'statistics', 'use': 'Matplotlib styles for scientific publication contexts.', 'tags': 'publication style journal', 'official_url': 'https://github.com/garrettj403/SciencePlots'}, {'name': 'statannotations', 'language': 'Python', 'category': 'statistics', 'use': 'Statistical annotations for seaborn and Matplotlib plots.', 'tags': 'significance annotation box', 'official_url': 'https://github.com/trevismd/statannotations'}, {'name': 'GeoPandas', 'language': 'Python', 'category': 'specialized', 'use': 'Geospatial vector data analysis and mapping.', 'tags': 'map geometry spatial', 'official_url': 'https://geopandas.org/en/stable/docs/user_guide/mapping.html'}, {'name': 'Cartopy', 'language': 'Python', 'category': 'specialized', 'use': 'Projected maps and geospatial data transformations.', 'tags': 'map projection geospatial', 'official_url': 'https://cartopy.readthedocs.io/stable/gallery/index.html'}, {'name': 'Folium', 'language': 'Python', 'category': 'specialized', 'use': 'Interactive Leaflet maps from Python.', 'tags': 'interactive map leaflet', 'official_url': 'https://python-visualization.github.io/folium/latest/getting_started.html'}, {'name': 'pydeck', 'language': 'Python', 'category': 'specialized', 'use': 'Large-scale layered geospatial visualizations with deck.gl.', 'tags': 'map layers gpu spatial', 'official_url': 'https://deckgl.readthedocs.io/en/latest/gallery/index.html'}, {'name': 'Datashader', 'language': 'Python', 'category': 'specialized', 'use': 'Rasterize and aggregate massive point or line datasets.', 'tags': 'large data density raster', 'official_url': 'https://datashader.org/user_guide/index.html'}, {'name': 'NetworkX', 'language': 'Python', 'category': 'specialized', 'use': 'Network analysis, layouts, and graph drawing.', 'tags': 'network graph topology', 'official_url': 'https://networkx.org/documentation/stable/auto_examples/index.html'}, {'name': 'PyVis', 'language': 'Python', 'category': 'specialized', 'use': 'Interactive browser-based network visualization.', 'tags': 'network interactive browser', 'official_url': 'https://pyvis.readthedocs.io/en/latest/'}, {'name': 'OSMnx', 'language': 'Python', 'category': 'specialized', 'use': 'Download, analyze, and visualize street networks.', 'tags': 'network street map osm', 'official_url': 'https://osmnx.readthedocs.io/en/stable/'}, {'name': 'contextily', 'language': 'Python', 'category': 'specialized', 'use': 'Add web basemaps to geospatial plots.', 'tags': 'basemap tiles spatial', 'official_url': 'https://contextily.readthedocs.io/en/latest/intro_guide.html'}, {'name': 'GeoViews', 'language': 'Python', 'category': 'specialized', 'use': 'Geographic visualization built on HoloViews.', 'tags': 'map geographic interactive', 'official_url': 'https://geoviews.org/gallery/index.html'}, {'name': 'Graphviz', 'language': 'Python', 'category': 'specialized', 'use': 'Graph and hierarchy layout through the Graphviz engine.', 'tags': 'network tree layout', 'official_url': 'https://graphviz.readthedocs.io/en/stable/examples.html'}, {'name': 'pyCirclize', 'language': 'Python', 'category': 'specialized', 'use': 'Circular genome, chord, and sector visualizations.', 'tags': 'circular chord polar', 'official_url': 'https://moshi4.github.io/pyCirclize/'}, {'name': 'ggplot2', 'language': 'R', 'category': 'grammar', 'use': 'Layered grammar of graphics for analytical and publication figures.', 'tags': 'grammar layers facets', 'official_url': 'https://ggplot2.tidyverse.org/'}, {'name': 'plotly for R', 'language': 'R', 'category': 'grammar', 'use': 'Interactive Plotly figures and ggplotly conversion.', 'tags': 'interactive hover web', 'official_url': 'https://plotly.com/r/'}, {'name': 'Shiny', 'language': 'R', 'category': 'grammar', 'use': 'Reactive interactive applications and dashboards.', 'tags': 'app reactive dashboard', 'official_url': 'https://shiny.posit.co/r/gallery/'}, {'name': 'highcharter', 'language': 'R', 'category': 'grammar', 'use': 'Interactive Highcharts visualizations from R.', 'tags': 'interactive web time series', 'official_url': 'https://jkunst.com/highcharter/'}, {'name': 'echarts4r', 'language': 'R', 'category': 'grammar', 'use': 'R interface to Apache ECharts with rich interaction.', 'tags': 'interactive echarts web', 'official_url': 'https://echarts4r.john-coene.com/'}, {'name': 'ggiraph', 'language': 'R', 'category': 'grammar', 'use': 'Interactive SVG output for ggplot2 graphics.', 'tags': 'ggplot interactive svg', 'official_url': 'https://davidgohel.github.io/ggiraph/'}, {'name': 'lattice', 'language': 'R', 'category': 'grammar', 'use': 'Trellis graphics for multivariable conditioning.', 'tags': 'facets conditioning multivariate', 'official_url': 'https://lattice.r-forge.r-project.org/'}, {'name': 'r2d3', 'language': 'R', 'category': 'grammar', 'use': 'Build custom D3 visualizations from R.', 'tags': 'd3 custom interactive', 'official_url': 'https://rstudio.github.io/r2d3/'}, {'name': 'htmlwidgets', 'language': 'R', 'category': 'grammar', 'use': 'Framework connecting JavaScript visualization libraries to R.', 'tags': 'html javascript interactive', 'official_url': 'https://www.htmlwidgets.org/showcase_leaflet.html'}, {'name': 'vegawidget', 'language': 'R', 'category': 'grammar', 'use': 'Render and compose Vega and Vega-Lite specifications.', 'tags': 'vega declarative grammar', 'official_url': 'https://vegawidget.github.io/vegawidget/'}, {'name': 'dygraphs', 'language': 'R', 'category': 'grammar', 'use': 'Interactive time-series charts for R and Shiny.', 'tags': 'interactive time series', 'official_url': 'https://rstudio.github.io/dygraphs/'}, {'name': 'reactable', 'language': 'R', 'category': 'grammar', 'use': 'Interactive data tables with sorting and grouping.', 'tags': 'table interactive arranged', 'official_url': 'https://glin.github.io/reactable/articles/examples.html'}, {'name': 'patchwork', 'language': 'R', 'category': 'statistics', 'use': 'Compose multiple ggplot2 figures with a layout grammar.', 'tags': 'composition panels publication', 'official_url': 'https://patchwork.data-imaginist.com/'}, {'name': 'cowplot', 'language': 'R', 'category': 'statistics', 'use': 'Align, arrange, annotate, and theme publication plots.', 'tags': 'publication arrange annotation', 'official_url': 'https://wilkelab.org/cowplot/'}, {'name': 'ggridges', 'language': 'R', 'category': 'statistics', 'use': 'Ridgeline density plots for grouped distributions.', 'tags': 'ridgeline distribution density', 'official_url': 'https://wilkelab.org/ggridges/'}, {'name': 'ggdist', 'language': 'R', 'category': 'statistics', 'use': 'Uncertainty and distribution visualization for ggplot2.', 'tags': 'uncertainty interval distribution', 'official_url': 'https://mjskay.github.io/ggdist/'}, {'name': 'ggbeeswarm', 'language': 'R', 'category': 'statistics', 'use': 'Non-overlapping point plots for distributions.', 'tags': 'beeswarm points distribution', 'official_url': 'https://eclarke.github.io/ggbeeswarm/'}, {'name': 'GGally', 'language': 'R', 'category': 'statistics', 'use': 'Pairs plots and extensions to ggplot2.', 'tags': 'pairplot correlation multivariate', 'official_url': 'https://ggobi.github.io/ggally/'}, {'name': 'ggstatsplot', 'language': 'R', 'category': 'statistics', 'use': 'ggplot2 figures integrated with statistical details.', 'tags': 'statistics annotation inference', 'official_url': 'https://indrajeetpatil.github.io/ggstatsplot/'}, {'name': 'forestplot', 'language': 'R', 'category': 'statistics', 'use': 'Customizable forest plots and confidence intervals.', 'tags': 'forest interval meta analysis', 'official_url': 'https://cran.r-project.org/package=forestplot'}, {'name': 'survminer', 'language': 'R', 'category': 'statistics', 'use': 'Publication-ready survival curves and diagnostics.', 'tags': 'survival model curve', 'official_url': 'https://rpkgs.datanovia.com/survminer/'}, {'name': 'ggforce', 'language': 'R', 'category': 'statistics', 'use': 'Geometric extensions, facets, and annotations for ggplot2.', 'tags': 'geometry facet annotation', 'official_url': 'https://ggforce.data-imaginist.com/'}, {'name': 'ggthemes', 'language': 'R', 'category': 'statistics', 'use': 'Additional complete themes, scales, and palettes for ggplot2.', 'tags': 'theme publication palette', 'official_url': 'https://jrnold.github.io/ggthemes/'}, {'name': 'ggtext', 'language': 'R', 'category': 'statistics', 'use': 'Rich text rendering inside ggplot2 figures.', 'tags': 'typography annotation publication', 'official_url': 'https://wilkelab.org/ggtext/'}, {'name': 'sf', 'language': 'R', 'category': 'specialized', 'use': 'Simple-features vector data operations and mapping.', 'tags': 'map geometry spatial', 'official_url': 'https://r-spatial.github.io/sf/'}, {'name': 'terra', 'language': 'R', 'category': 'specialized', 'use': 'Spatial raster and vector analysis with plotting support.', 'tags': 'raster map spatial', 'official_url': 'https://rspatial.github.io/terra/'}, {'name': 'tmap', 'language': 'R', 'category': 'specialized', 'use': 'Thematic static and interactive maps.', 'tags': 'thematic map interactive', 'official_url': 'https://r-tmap.github.io/tmap/'}, {'name': 'leaflet', 'language': 'R', 'category': 'specialized', 'use': 'Interactive Leaflet maps from R and Shiny.', 'tags': 'interactive map tiles', 'official_url': 'https://rstudio.github.io/leaflet/'}, {'name': 'mapview', 'language': 'R', 'category': 'specialized', 'use': 'Rapid interactive viewing of spatial objects.', 'tags': 'interactive map exploratory', 'official_url': 'https://r-spatial.github.io/mapview/'}, {'name': 'ggraph', 'language': 'R', 'category': 'specialized', 'use': 'Grammar-of-graphics approach to networks and trees.', 'tags': 'network tree grammar', 'official_url': 'https://ggraph.data-imaginist.com/'}, {'name': 'igraph', 'language': 'R', 'category': 'specialized', 'use': 'Network analysis, layout, and plotting.', 'tags': 'network topology graph', 'official_url': 'https://r.igraph.org/'}, {'name': 'ComplexHeatmap', 'language': 'R', 'category': 'specialized', 'use': 'Highly composable heatmaps with annotations.', 'tags': 'heatmap matrix annotation', 'official_url': 'https://jokergoo.github.io/ComplexHeatmap-reference/book/'}, {'name': 'ggalluvial', 'language': 'R', 'category': 'specialized', 'use': 'Alluvial plots for categorical flows in ggplot2.', 'tags': 'alluvial sankey flow', 'official_url': 'https://corybrunson.github.io/ggalluvial/'}, {'name': 'circlize', 'language': 'R', 'category': 'specialized', 'use': 'Circular, chord, genomic, and sector visualizations.', 'tags': 'circular chord polar', 'official_url': 'https://jokergoo.github.io/circlize_book/book/'}, {'name': 'networkD3', 'language': 'R', 'category': 'specialized', 'use': 'Interactive D3 networks, trees, and Sankey diagrams.', 'tags': 'network sankey interactive', 'official_url': 'https://christophergandrud.github.io/networkD3/'}, {'name': 'treemapify', 'language': 'R', 'category': 'specialized', 'use': 'Treemaps in the ggplot2 ecosystem.', 'tags': 'treemap hierarchy area', 'official_url': 'https://wilkox.org/treemapify/'}])
selected_tools = tool_catalog.query("language == @LANGUAGE").copy()
if TOOL_CATEGORY != "All":
selected_tools = selected_tools.query("category == @TOOL_CATEGORY")
if TOOL_SEARCH.strip():
query = TOOL_SEARCH.lower().strip()
selected_tools = selected_tools[
selected_tools.apply(
lambda row: query in " ".join(map(str, row)).lower(), axis=1
)
]
display(selected_tools[["name", "language", "category", "use", "official_url"]].style.hide(axis="index"))
print(
"Complete atlas:",
len(tool_catalog), "tools ·",
(tool_catalog["language"] == "Python").sum(), "Python ·",
(tool_catalog["language"] == "R").sum(), "R",
)
3.5 Open science and data governance—three paper-anchored questions¶
| Question | Inspect | Course anchors |
|---|---|---|
| Can we trace and reproduce it? | Producer, original source, revision, schema, license, selection, transformation | FAIR — Wilkinson et al. (2016); Croissant — Akhtar et al. (2024) |
| Is it fit for this claim? | Coverage, measurement, missingness, uncertainty, visual implication, non-inference | Lan & Liu (2024); Ziman et al. (2026) |
| Who benefits, controls, and bears risk? | Authority, consent, context, correction, withdrawal, intended use | CARE + FAIR — Carroll et al. (2021) |
# @title 3.6 · Model a citation-based governance critique
governance_audit = pd.DataFrame([
[
"Observed",
"The Dataset Card identifies a public Parquet conversion and CC BY 4.0; the upstream dump documents the positional schema.",
"Dataset Card + GeoNames export",
"Record the exact SHA, field mapping, IDs, and transformations.",
],
[
"Inferred",
"The rows make both locations geographically relevant but cannot establish cultural equivalence or environmental change.",
"Lan & Liu (2024); Ziman et al. (2026)",
"Keep unsupported claims outside the visual conclusion.",
],
[
"Open question",
"The source does not establish whether local communities defined the comparison, benefits, or acceptable reuse.",
"Carroll et al. (2021)",
"Validate purpose, authority, and possible harm with affected communities.",
],
], columns=["evidence status", "claim", "basis", "design consequence"])
display(governance_audit.style.hide(axis="index"))
# @title 3.7 · Add evidence only to repair a named gap
additional_evidence = pd.DataFrame([
[
"Cultural authority",
"UNESCO Venice and its Lagoon",
"https://whc.unesco.org/en/list/394/",
"Designation context; not everyday community experience or Kunshan priorities",
],
[
"Environmental change",
"JRC Global Surface Water",
"https://global-surface-water.appspot.com/download",
"Water occurrence/change; not cause, quality, or cultural meaning",
],
[
"Spatial context",
"OpenStreetMap",
"https://www.openstreetmap.org/copyright",
"Waterways and places; contributor coverage is uneven and requires verification",
],
], columns=["gap repaired", "candidate source", "direct URL", "evidence boundary"])
display(additional_evidence.style.hide(axis="index"))
# @title 3.8 · Replace the instructor case and generate three working prompts
student_project = {
"title": "REPLACE",
"domain_question": "REPLACE",
"intended_community": "REPLACE",
"dataset_url": "REPLACE",
"dataset_and_attribute_types": "REPLACE",
"actions_and_targets": "REPLACE",
"idiom_decision": "REPLACE",
"algorithm_and_tools": "REPLACE",
"evidence_boundary": "REPLACE",
"additional_source": "REPLACE",
}
generation_prompt = f"""
Create one complete index.html for {student_project['title']}.
Domain question: {student_project['domain_question']}
Community: {student_project['intended_community']}
Dataset: {student_project['dataset_url']}
Data/attributes: {student_project['dataset_and_attribute_types']}
Tasks as action + target: {student_project['actions_and_targets']}
Use verified source rows, visible attribution, accessible browser-compatible code,
explicit loading/error states, and no unsupported claims.
"""
deployment_prompt = """
Convert the attached index.html into a zero-build Hugging Face Static Space.
Preserve interactions, use HTTPS resources, add visible provenance and limitations,
and return complete index.html plus README.md.
"""
critique_prompt = f"""
Critique the deployed Space at four levels: domain; data/task; idiom; algorithm.
Apply the assigned FAIR, CARE, Croissant, and visualization papers to direct source
evidence. Separate Observed, Inferred, and Open Question claims. Redesign using:
idiom={student_project['idiom_decision']};
algorithm={student_project['algorithm_and_tools']};
evidence boundary={student_project['evidence_boundary']};
additional source={student_project['additional_source']}.
"""
if any(value == "REPLACE" for value in student_project.values()):
print("Instructor example mode · replace every REPLACE field before using these prompts.")
display(HTML("<h3>Generation prompt</h3><pre>" + generation_prompt.strip() + "</pre>"))
display(HTML("<h3>Static-Space prompt</h3><pre>" + deployment_prompt.strip() + "</pre>"))
display(HTML("<h3>Critique/redesign prompt</h3><pre>" + critique_prompt.strip() + "</pre>"))
# @title 3.9 · Export the reproducibility bundle
attribute_contract.to_csv(OUTPUT_DIR / "data_attribute_contract.csv", index=False)
task_contract.to_csv(OUTPUT_DIR / "action_target_contract.csv", index=False)
evidence_boundary.to_csv(OUTPUT_DIR / "evidence_boundary.csv", index=False)
pipeline.to_csv(OUTPUT_DIR / "four_level_pipeline.csv", index=False)
idiom_catalog.to_csv(OUTPUT_DIR / "idiom_catalog_38.csv", index=False)
tool_catalog.to_csv(OUTPUT_DIR / "python_r_tool_atlas_72.csv", index=False)
governance_audit.to_csv(OUTPUT_DIR / "governance_audit.csv", index=False)
additional_evidence.to_csv(OUTPUT_DIR / "additional_evidence_plan.csv", index=False)
(OUTPUT_DIR / "student_project.json").write_text(
json.dumps(student_project, indent=2, ensure_ascii=False), encoding="utf-8"
)
bundle = shutil.make_archive("INFOSCI301_INFOVIS_Redesign_Colab_Outputs", "zip", OUTPUT_DIR)
expected = {
"geonames_water_towns.csv",
"jiangnan_map.html",
"venice_map.html",
"population_comparison.png",
"population_comparison.svg",
"population_comparison.pdf",
"four_level_pipeline.csv",
"idiom_catalog_38.csv",
"python_r_tool_atlas_72.csv",
"governance_audit.csv",
}
present = {path.name for path in OUTPUT_DIR.iterdir()}
assert expected <= present, sorted(expected - present)
print("Export complete:", Path(bundle).resolve())
print("Verified:", len(records), "source rows ·", len(idiom_catalog), "idioms ·", len(tool_catalog), "tools")
Finish with three judgments¶
- Claim: What can the current evidence support—and what must remain explicitly unclaimed?
- Choice: Which human decision changed the data, idiom, algorithm, or interaction after critique?
- Consequence: Whose understanding, authority, benefit, or risk should the next test evaluate?
References¶
Visualization: Munzner (2014), book companion; Pu & Kay (2023), CHI paper; Lan & Liu (2024), TVCG paper; Ziman et al. (2026), TVCG paper.
Open science and governance: Wilkinson et al. (2016), FAIR; Carroll et al. (2021), CARE + FAIR; Akhtar et al. (2024), Croissant.
Data and implementation: do-me/Geonames; GeoNames export; Folium documentation; OpenStreetMap tile policy.