Center for Geographic Analysis | Harvard University
2025-11-07
Spatial Data, Data Science, and AI … and red herrings
Red herring: a distraction from the real problem(s) at hand
We live in exciting times for data and methods.
We also live in “interesting” times for applications.
And complicated disciplinary times.
And so much more than “new and emerging forms of data”
Historical – maps, diaries, inventories, censuses, surveys, archeological sites
Big and smart – sensors, social media, internet, mobility, street view
Government and administrative – census, medical, vital records, administrative
Linked – joining or nesting information across locations, time, or individual
From space and from the air – satellite imagery, aerial photography, drones, LiDAR
Not just technical – also social
Requires investment and maintenance
Crumbles with lack of attention
Easily taken for granted
“Traditional” methods to ascertain relationships and associations across variables and places, collapse data dimensionality, compare groups and outcomes
Data science and visualization
Machine learning
AI: computing and analytical innovations that facilitate data discovery and manipulation, text analysis, feature extraction, data creation, and analytics at scale
(and challenges)
To tackle wicked social, environmental, and health challenges
Increase understanding of the world around us
Do really interesting and innovative research
And train the next generation to do even better
But…
As social scientists we can occasionally get distracted by the shiny data and methods objects. The attraction of the novel.
The ways in which data and methods distract us from the real problem of improving well-being and reducing spatial inequalities.
This is the tortured geographer part of the talk.
(which may often be “AI”!)
Examples: mobile phone data, social media, even satellite imagery
Example: socio-economic and demographic characteristics for mobility data
Thinking about how emerging technologies intersect with the spatial demography of cities to exacerbate, reproduce, and generate inequalities across areas or groups.
How can we support informed and equitable decision-making around sensor placement, especially what criteria ideal networks might satisfy and the inevitable trade-offs involved?
What is the ”best” allocation of n sensors, given a particular goal?
Decision support tools that visualize options and trade-offs
Coverage for older residents (>65)
Coverage for place-of-work population
But should make us think:
how are our (air quality) data produced?
how can we do better?
Thinking about how health, climate change, and transportation intersect
Only 4 out of 11 London underground lines have air conditioning systems
The average summer temperature in London is expected to increase by 2.7 degrees Celsius by the 2050s
The probability of heatwaves could also increase five-fold and they’re expected to occur every other year
By 2070, the mean maximum air temperature in the UK in August is projected to increase by up to 6 °C in summer compared to 2018
How can we estimate current and future heat exposure on the Tube and who (where) is most affected?
Who’s doing the travelling? (demographic and health characteristics)
Where are they going and how do they get there? (origin and destination flows)
What’s the temperature at the station and on board? (estimated from known station-surface differentials)
Synthetic Population Catalyst (SPC)–A synthetic population that simulates individual-level travel behavior (home and workplace), travel mode, demographic and socio-economic characteristics that allow us to explore heat vulnerability at the individual level
Tube operation timetable–For travel times and route estimation, we use the timetable provided by Transport for London (TfL) APIs. Provides accurate estimation whether travelers for each OD-pair will take air-conditioned Tube lines
Clim-recal–Estimates weather and heat wave days in the past and future on a daily basis in 2.2 km*2.2 km cells covering the entire UK. Local variation in the dataset is used to estimate heat exposure more accurately.
Extreme Degree Minutes
More vulnerable populations are (slightly more) impacted
Non-white ethnicities also (slightly more) impacted
Unable to really look at age, which is a big deal
The inequality of heat exposure risk is significant in spatial terms
Trickiness of estimation but lots of useful data that can be brought to bear
An exemplary example of where geographers have a lot to contribute
But also some of this should be being measured directly!
Making satellite imagery data usable, useful, and used in the social sciences, health, and policy
We’ve now got the compute, methods, and sensor quality for satellite imagery to be a game-changer for social science and health research and policy making
1. Imagery innovation–research-ready imagery-based data products, building off and developing innovative computing and AI methods that facilitate efficient automated workflows for measures and indicators
2. Data for all–data distribution channels that meet researchers and policymakers where they are
3. Capability and community–building capacity for understanding and working with imagery and imagery-derived data
AI and novel data are true game changers
Novel insights and applications
Data (and methods to exploit that data) that give us completely new views into human behavior and preferences
At scale
Useful but also fun
Industry- versus publicly-owned data
Lack of investment into novel forms of public data
Preoccupation with fast, new, shiny instead of slow, theoretically grounded and perhaps actually impactful
Bandwagons (AI or otherwise)
Spillover effects: on publication, hiring, external research funding…
(Not the other way around)
2025 Malcolm Comeaux Lecture | Arizona State University | November 7, 2025