Lecture 3

Solar Adoption Data: Measurement, Theory, and Research Design

Byeong-Hak Choe

SUNY Geneseo

August 31, 2026

🎯 Today we turn one distinction into a research plan

By the end of class, you should be able to:

  1. distinguish technical potential from observed adoption;
  2. connect environmental-economic mechanisms to variables;
  3. audit a building-level snapshot without changing its meaning; and
  4. choose one bounded contribution to the shared project.

Activate what you already know: NPV, externalities, subsidies, credit constraints, dplyr, and careful denominators.

🏠 Two sunny roofs can produce two different choices

Same sunlight. Similar roof. What could explain the difference?

Suitable roof A

Similar sunshine · similar usable area

Installs solar

Suitable roof B

Similar sunshine · similar usable area

Does not install

Returns · finance · information · tenure · installers · policy · preferences · risk

🔐 Research design includes permissions and credentials

01 · Secure

  • No API key in source code.
  • Restrict replacement keys.
  • Monitor usage.

02 · Verify

  • Review current storage rules.
  • Review publication and attribution.
  • Confirm permitted use before class access.

03 · Minimize

  • No raw coordinates or building IDs for students.
  • No Google content in course ML.

🪜 Potential is only the first rung of the outcome ladder

Solar resourcesunshine
Technical potentialroof + layout
Economic potentialNPV
Market potentialconstraints
Realized adoptioninstallation
Realized outcomesgeneration + bills + emissions

Each arrow adds actors, choices, data, and uncertainty. Later outcomes cannot be inferred automatically from earlier rungs.

🧭 Which gap? Technical potential is not the benchmark

01 · TECHNICALWhat is physically possible?

roof · sunshine · modeled layouts

Google snapshot
02 · PRIVATEWhat is worthwhile for this decision-maker?

tariff · price · finance · tenure · risk · hassle

not observed
04 · OBSERVEDWhat actually happened?

installation · timing · realized output

needs adoption data
Deployment gaptechnical ↔︎ observed
Private investment gapprivate optimum ↔︎ observed
Social welfare gapsocial optimum ↔︎ observed

The Google snapshot populates one benchmark; it cannot diagnose an inefficient gap by itself.

💵 Adoption occurs when expected private value clears private costs

ADOPTi = 1 if
expected bill savings + incentives + other benefits
> capital + finance + transaction + perceived-risk costs

Heterogeneity changes every term.

Time

discount rate · waiting · policy risk

Money

cash · credit · tax appetite · financing

Friction

search · permits · disruption · quality risk

🧩 Four explanations for the same nonadoption

MODEL / MEASUREMENTThe projected return is wrong.

Generic tariff · old imagery · omitted roof repair

Compare models with quotes, production, and updated roofs.
RATIONAL HETEROGENEITYWaiting is privately reasonable.

Moving risk · replacement timing · uncertainty · option value

Observe household circumstances and adoption timing.
MARKET FAILUREA correctable wedge blocks adoption.

Information · credit · split incentives · spillovers

Use policy, credit, tenure, or information variation.
BEHAVIORAL FRICTIONThe decision process is distorted.

Inattention · present bias · inertia · complexity

Test targeted information, simplification, or reminders.

Policy follows the diagnosed mechanism—not the size of the raw technical-potential gap.

⚖️ Private incentives and social value need not align

PRIVATE NPV ≠ SOCIAL NPV

Household / firm

  • avoided retail bills;
  • export compensation;
  • tax credits and rebates; and
  • resilience and preferences.

Society

  • avoided marginal generation;
  • climate and local pollution;
  • network and integration costs; and
  • learning and information spillovers.

Policy may correct an externality—but tariff and subsidy design also determine who can respond and who pays.

🌐 Opportunity, ability, and access differ

Opportunity

sunshine · roof geometry · ownership

Ability

income · liquidity · credit · tax appetite

Access

information · installer quotes · programs

technical potential → economic feasibility → realized adoption

Equity analysis asks where groups fall out of the funnel—and why.

🔬 A theory becomes testable when it implies a comparison

Mechanism Prediction More credible comparison
Up-front rebate Adoption rises when immediate cost falls Before/after a discrete policy change; nearby control markets
Peer learning Prior nearby installs raise later adoption Exposure over time; rule out shared neighborhood demand
Installer access Fewer viable quotes suppress adoption Quote availability holding measured demand and site traits fixed
Credit / liquidity High-NPV sites still fail to convert Comparable potential across financing access or tenure

Prediction + comparison + assumptions = an empirical design. A regression alone is not the design.

📚 Policy evidence: incentives work—but incidence matters

Hughes & Podolefsky (2015)
Rebate discontinuity / quasi-experiment
A rebate change increased installations; substantial inframarginal transfers remain possible.
De Groote & Verboven (2019)
Dynamic structural model
Households heavily discount future subsidies; up-front support can be more effective.
Borenstein (2017)
Tariff and subsidy accounting
Retail rate design can shape private solar value as much as the export credit.
Sexton et al. (2021)
Modeled heterogeneous benefits
Uniform subsidies can be poorly aligned with avoided pollution benefits.

Policy can raise adoption and still misallocate benefits or transfers.

🧑‍🤝‍🧑 Diffusion and equity: clusters are clues, not causal answers

Diffusion / market frictions

Equity / participation

Spatial clustering may reflect peers, sorting, common policy, or shared installers.

🧾 An evidence matrix disciplines literature claims

Question Design / data Estimand Limit
Do rebates change installs? Discrete rebate change Local response to the policy External validity; inframarginal recipients
Do peers affect adoption? Prior nearby installs over time Effect of additional prior exposure Shared shocks and sorting
Who receives quotes? Installer-platform inquiries Conditional quote gap Selection into requesting quotes
Who adopts? Installed-system + income data Descriptive participation gap Not a causal mechanism

Pair prompt (2 minutes): Choose one result. What can it support—and what would overclaim?

🌞 Umbrella question: where does rooftop potential become adoption?

How do technical potential, household and neighborhood conditions, and policy or market access jointly shape realized distributed-solar deployment?

01 · Measure

Where is modeled rooftop potential—and how reliable is the match?

How does open deployment data align with geography and time?

03 · Explain

Which mechanisms predict conversion, gaps, and heterogeneity?

First milestone: a defensible descriptive baseline—not a causal verdict.

🔽 10,000 queries narrow to 4,829 valid matches

10,000candidate points
8,311API returned a building
4,829≤ 30 m valid building match
4,791valid + potential values

48.29% of input points become valid building matches.

📍 “Found” means nearest returned building; distance decides validity

≤ 30 m
> 30 m
input point
returned building
12.9 m — valid
returned building
> 30 m — invalid

3,482 points returned a building but failed the 30 m rule. Median distance among valid matches: 12.89 m.

📊 Coverage varies across the nine sampled metros

FOUNDVALID
New York City
84.6%
55.2%
Rochester
92.6%
50.4%
Syracuse
90.8%
42.5%
Albany
87.3%
37.7%
Buffalo
63.9%
36.1%
Utica–Rome
80.0%
35.0%
Binghamton
91.0%
33.7%
Ithaca
86.4%
33.2%
Geneseo
44.4%
9.2%

Percent of sampled input points · denominator differs by metro

Coverage and matching diagnostics—not solar-adoption rates.

🗺️ Statewide rooftop suitability needs a different benchmark

Map of New York counties shaded by the estimated share of small buildings with a PV-suitable roof plane. Gold dots mark the nine Google sample metros.

STATEWIDE MODEL62 counties

NREL ZIP estimates aggregated with modeled small-building counts.

PROJECT SNAPSHOT9 sampled metros

Newer building detail—but not a statewide building sample.

Estimated share of small buildings with a PV-suitable roof plane · NREL 2016 model

This is a physical benchmark—not installed solar, private NPV, or an adoption rate.

📏 The building distribution has a long right tail

Maximum modeled panel count among valid buildings with potential

MEDIAN43
MEAN131
MAX20,567
1101001k10k100k

One typical building is much smaller than the mean. Report medians and IQRs before averages—and inspect the extreme structures.

🗂️ Multiple table grains—and one export trap

Table One row represents… Rows
candidates one input point + nearest returned building 10,000
valid_buildings one input point passing the 30 m rule 4,829
roof_segments one modeled roof segment 78,351
panels one algorithmic panel placement 2,310,912
panel_configs one modeled system configuration 682,527
financial_analyses one hypothetical monthly-bill scenario 189,267
Export trap: Child files include found-but-invalid matches. For panels, 1,682,259 of 2,310,912 rows are attached to invalid parents.
panels |>
  semi_join(valid_buildings,
            by = "input_id")

👥 Work in pairs: one driver, one denominator auditor

Driver

  • reads the file;
  • states the row grain;
  • writes one dplyr pipeline; and
  • narrates each transformation.

Denominator auditor

  • names the denominator;
  • checks unique keys;
  • predicts join multiplication; and
  • challenges one causal word.

Swap roles after the first result. Your deliverable is one table and one cautious sentence.

💻 Beginner audit: valid-match rates by metro

candidate_file <- file.path(
  "/Volumes/Extreme SSD",
  "solar-adoption", "data",
  paste0("solar_building_insights_",
         "all_candidates_2026-05-28.csv"))
candidates <- read_csv(candidate_file)

metro_audit <- candidates |>
  group_by(metro) |>
  summarise(
    input_n = n(),
    found_n = sum(found),
    valid_n = sum(valid_building_match),
    valid_rate = valid_n / input_n
  ) |>
  arrange(desc(valid_rate))

Before running

  1. One output row represents…?
  2. What is the denominator of valid_rate?
  3. Why is the mean of the TRUE/FALSE match flag equivalent here?
  4. What claim can this table support?
Predict → run → inspect → explain
THIS AUDIT · FACULTY SSD/Volumes/Extreme SSD/solar-adoption/data/
CLASSWORK 2 · PUBLICnyserda_solar_county_year_2026-06-30.csv

🧪 Three audit findings change the analysis before modeling begins

SELECTION67.45%of all valid matches come from New York City
GEOGRAPHY DRIFT531valid returned buildings lie outside New York State
TIME MISMATCH2012–2024imagery dates underlie a May 2026 snapshot

Sampling, boundaries, and vintage belong in every estimate—not only in a limitations paragraph.

✍️ Match claims to the strength of the data

Defensible now

“Among sampled points in the NYC metro, 55.2% had a returned building within 30 meters.”

“Modeled panel capacity is highly right-skewed among valid matches.”

Not supported

“55.2% of NYC buildings adopted solar.”

“Low-income households have less rooftop potential.”

“Peer effects caused the observed clusters.”

Name the sample · unit · measure · denominator · time · uncertainty.

🧱 Work begins with an open-data backbone

NYSERDA

observed distributed-solar projects

deployment outcome

ACS 5-year

income · tenure · housing · demographics

community context

TIGER/Line

tracts · counties · boundaries

spatial join frame

Faculty-only checkpoint: use the retained Google potential snapshot only after institutional review of terms, retention, attribution, and student access.

Open outcome + open context first; restricted potential may be joined later.

🛤️ A five-step path into the shared project

1
REPRODUCErun an existing audit
2
AUDITdocument grain, keys, and limits
3
PROPOSEone bounded issue or question
4
BRANCHisolated QMD, script, or note
5
REVIEWpeer audit, then faculty merge

Rotating roles: literature / theory · provenance / QA · open adoption + ACS · visualization / spatial · reproducibility

No billable API calls · no credentials · no raw property coordinates · no restricted-data ML

✅ Choose your first contribution

  1. Write one researchable question.
  2. Name its outcome and denominator.
  3. Choose one mechanism and one comparison.
  4. State one claim the available data cannot support.

Potential is measured. Adoption is chosen. Research explains the gap.