The projected return is wrong.
Old imagery, a generic tariff, roof-repair needs, or optimistic output can make modeled value too high.
Look for: model–quote or prediction–production errors that vary systematically across places.Use economics to explain the steps between a technically suitable roof and a completed installation—then match each explanation to evidence that New York data can actually provide.
A modeled roof is an opportunity, not an installed system or an economic optimum.
Begin with county-year data; move to ZIP- or tract-time only when the underlying records permit it.
Call a pattern causal only when policy or program variation supplies a credible counterfactual.
The project’s umbrella question is:
How do technical potential, household and neighborhood conditions, and policy or market access jointly shape realized distributed-solar deployment in New York?
The outcome is not one jump from sunlight to solar panels. It is a sequence of increasingly demanding conditions.
Technical-potential studies estimate what suitable roofs could generate if used, independent of financial or social constraints (Gagnon et al. 2016). Google’s modeled roof geometry, sunshine, and panel layouts therefore populate an early rung of this ladder. They do not reveal whether a household installed panels, whether installation would be privately profitable, or whether it would maximize social welfare. See the Project Sunroof methodology and Solar API methodology for the construction of these measures.
The energy-efficiency-gap literature is useful here because it forces the analyst to name the benchmark before diagnosing “too little” adoption (Jaffe and Stavins 1994; Allcott and Greenstone 2012; Gerarden et al. 2017).
Google helps measure this.
Tariffs, financing, tenure, risk, and hassle are needed.
Grid effects, emissions, externalities, and real resource costs are needed.
NYSERDA records projects, timing, and capacity.
| Gap being described | Comparison | What it means | Can current data identify it? |
|---|---|---|---|
| Deployment gap | Technical potential ↔︎ observed deployment | Some modeled opportunity is not realized | Descriptively, with careful aggregation |
| Private investment gap | Private optimum ↔︎ observed deployment | Privately worthwhile investment did not occur | No—household-specific costs and benefits are missing |
| Social welfare gap | Social optimum ↔︎ observed deployment | Deployment differs from the welfare-maximizing level or location | No—marginal grid and environmental values are missing |
Interpretation rule: a large technical-potential gap is not, by itself, evidence of irrational households, a market failure, or a welfare-improving subsidy.
For household or firm i, a useful organizing inequality is:
ADOPTi = 1 if expected bill savings + incentives + other private benefits exceed capital + finance + transaction + perceived-risk costs
Technical output enters expected bill savings, but nearly every other term varies across people, buildings, utilities, programs, and time. Borenstein shows that retail tariff design and tax incentives can materially change private solar value (Borenstein 2017). De Groote and Verboven show why the timing of benefits matters: households in their setting heavily discounted future production subsidies, so equivalent support delivered up front could induce adoption at much lower public cost (De Groote and Verboven 2019).
Reviews of the energy-efficiency gap emphasize omitted costs, heterogeneous preferences, uncertainty, market failures, and behavioral mechanisms (Gillingham and Palmer 2014; Gerarden et al. 2017). The same low-adoption place can fit four different stories:
Old imagery, a generic tariff, roof-repair needs, or optimistic output can make modeled value too high.
Look for: model–quote or prediction–production errors that vary systematically across places.Moving risk, replacement timing, uncertainty, preferences, or option value can make nonadoption optimal.
Look for: adoption timing and household or building circumstances—not only average potential.Credit constraints, split incentives, imperfect information, or spillovers separate private choices from an appropriate benchmark.
Look for: exogenous changes in finance, tenure, information, or incentives.Inattention, present bias, inertia, and complexity may suppress action even when returns are understood correctly.
Look for: randomized simplification, reminders, salience, or information treatments.The weatherization experiment by Fowlie, Greenstone, and Wolfram is the cautionary benchmark: realized savings were far below engineering projections, so apparently attractive modeled investments did not deliver the forecast returns (Fowlie et al. 2018). This does not prove that Google Solar overstates rooftop potential. It tells us to validate model-based production or value before treating a prediction as an economic return.
The unsigned “Energy Efficiency Gap” course-note excerpt in the reference folder is useful for teaching behavioral policy debates, but it is a synthesis without identifiable publication metadata. Use the original studies—not the excerpt—as research citations.
Solar incentives can increase installations, but an adoption effect alone does not tell us whether a policy is well targeted or socially efficient.
A California rebate increase raised installations by about 10%; their broader model also implies sizable inframarginal transfers (Hughes and Podolefsky 2015).
NY use: exploit discrete NY-Sun incentive changes, with a comparison market and technical-opportunity controls.Property-assessed clean-energy financing increased solar investment in the studied California setting (Kirkpatrick and Bennear 2014).
NY use: compare places before and after financing access, if rollout dates and eligibility are recoverable.A dynamic model finds strong discounting of future benefits and much greater cost-effectiveness from up-front support (De Groote and Verboven 2019).
NY use: model adoption timing around anticipated incentive-block changes; do not infer discounting from dates alone.California’s tiered tariffs created a private solar incentive nearly as large as the federal tax credit in the study period (Borenstein 2017).
NY use: add utility territory, historical retail rates, and compensation rules before interpreting private returns.Spatial clustering is consistent with peer learning, but it can also reflect common housing, sorting, shared policies, or the same installers. Bollinger and Gillingham estimate peer effects using adoption histories and a design intended to separate prior exposure from shared neighborhood demand (Bollinger and Gillingham 2012). Solarize campaigns caused additional installations and lower prices, consistent with social learning and lower customer-acquisition costs, but the study does not isolate a pure information channel (Gillingham and Bollinger 2021).
Information also need not cross the entire adoption funnel. In a randomized experiment in India, an information tool improved knowledge and increased strong intent to adopt, yet the effect on actual adoption was statistically insignificant (Mahadevan et al. 2023). In U.S. platform data, prospective customers in lower-income tracts received fewer installer quotes conditional on measured factors; the observational design cannot reveal installer targeting independently of who enters the platform (O’Shaughnessy et al. 2021).
NYSERDA observes the end of this funnel for recorded projects. Google primarily observes the first stage. Neither source directly observes awareness, quotes, credit approval, intent, or decision authority.
Adopters remain higher-income than the overall population, although participation has broadened over time (Forrester et al. 2024). Tenure adds a separate barrier because a renter may live under a suitable roof without holding the investment decision (Best et al. 2023). Installer attention can make access unequal even after a prospective customer seeks a quote (O’Shaughnessy et al. 2021).
Sunter, Castellanos, and Kammen reported large racial and ethnic deployment disparities using Project Sunroof and ACS tract data (Sunter et al. 2019). Dokshin and Thiede’s reconstruction produced different national magnitudes and emphasized filtering, normalization, state heterogeneity, urban sample bias, and ecological inference (Dokshin and Thiede 2023). Treat this as a methodological replication dispute—not as proof that the broader inequity question is settled in either direction.
Ecological-inference rule: a tract’s median income, renter share, or racial composition describes the area. It does not identify the income, tenure, or race of the household that adopted solar.
Private value depends on avoided retail bills and transfers such as rebates or tax credits. Social value depends on real resource costs, displaced marginal generation, local pollution, climate damages, grid effects, and learning spillovers. These values vary over time and place (Borenstein 2012). Sexton and coauthors show that avoided pollution benefits can be poorly aligned with subsidy levels and estimate large gains from better spatial targeting (Sexton et al. 2021).
bill savings + export compensation + incentives + resilience/preferences − private costs
avoided generation + climate/local pollution + system effects − real resource and integration costs
NYSERDA plus Google can describe deployment and modeled annual production. They cannot, by themselves, value marginal emissions, congestion, capacity, or cost shifting. A welfare analysis is a later-stage project requiring grid and emissions data.
The table below treats each paper as design → estimand → limitation → project use rather than as a free-floating conclusion.
| Study | Context and design | What it credibly contributes | Transfer to the NY project |
|---|---|---|---|
| Fowlie, Greenstone & Wolfram (Fowlie et al. 2018) | Michigan weatherization; randomized encouragement plus quasi-experimental evidence | Validation can overturn engineering-return projections | Seek independent production, quotes, or roof assessments; do not transfer the weatherization return estimate to solar |
| Borenstein (Borenstein 2017) | California billing, adopter, tariff, and incentive accounting | Tariff design and incentives strongly shape private value; adopter composition matters | Add utility territory, tariffs, consumption proxies, and incentive rules |
| Hughes & Podolefsky (Hughes and Podolefsky 2015) | California rebate changes; quasi-experimental analysis and model counterfactual | Local adoption response plus evidence of inframarginal transfers | Separate a local policy response from a modeled no-program counterfactual |
| De Groote & Verboven (De Groote and Verboven 2019) | Flanders; dynamic structural adoption model | Timing and expected future benefits matter for subsidy design | Use program schedules and expectations; do not label timing patterns “present bias” without a model |
| Bollinger & Gillingham (Bollinger and Gillingham 2012) | California adoption histories; peer-effect design | Prior nearby installations can affect later adoption | Use ZIP-time histories, spatial lags, and tests against shared shocks; county-year data are too coarse for replication |
| Gillingham & Bollinger (Gillingham and Bollinger 2021) | Solarize campaigns; program evaluation | Campaigns raised installations and lowered prices | Recover treated municipalities, dates, and comparison places before making causal claims |
| O’Shaughnessy et al. (O’Shaughnessy et al. 2021) | EnergySage inquiries and quotes; observational platform data | Supply-side access differs by area income, conditional on measured factors | Developer concentration is a useful proxy, but NYSERDA cannot reveal rejected or missing quotes |
| Mahadevan, Meeks & Yamano (Mahadevan et al. 2023) | Cluster-randomized information intervention in India | Knowledge and intent can improve without a detectable installation response | If outreach is tested in NY, preregister installation—not only intent—as the primary outcome |
| Sunter et al.; Dokshin & Thiede (Sunter et al. 2019; Dokshin and Thiede 2023) | Project Sunroof + ACS tract comparisons and replication | Equity estimates are sensitive to sample, normalization, and regional heterogeneity | Report area-level disparities, weights, exclusions, and sensitivity analyses |
| Dokshin, Gherghina & Thiede (Dokshin et al. 2024) | Nearly all NY residential incentive installations, 2010–2020; tract panels | NY disparities changed over time and differed sharply by region | Establishes the closest baseline; new work should add later years, technical opportunity, market mechanisms, or a sharper design |
Two studies are especially close to this project. Araújo, Boucher, and Aphale analyze early clean-energy adopters in New York and find that income and home value are important correlates, while local patterns are more nuanced than a single statewide story (Araújo et al. 2019). Dokshin, Gherghina, and Thiede use geocoded NYSERDA records for 100,124 residential installations initiated from 2010–2020—95.6% of the state’s residential installations through 2020 in their data—to study tract-level disparities (Dokshin et al. 2024). They find that racial, income, and rural–urban gaps evolved differently over time and that New York City/Westchester, Long Island, and Upstate followed distinct trajectories.
That evidence raises the bar for a new contribution. A new project should not merely show that high-income places have more solar. It should add at least one of the following:
Use public NYSERDA records through 2026 and test whether earlier regional or income patterns persist.
Use vetted Google measures to distinguish low deployment from low sampled roof suitability—without calling the sample a statewide census.
Bring in incentive blocks, utility tariffs, Solarize timing, developers, or permitting to test a specific channel.
Show how geography, coverage, imagery vintage, eligible housing, and outcome normalization change the result.
| Source | Economic role | Useful measures | Appropriate join | Cannot establish |
|---|---|---|---|---|
| NYSERDA Statewide Distributed Solar Projects | Realized deployment | project timing, ZIP/county, capacity, estimated production, utility/developer fields in the full source | ZIP- or county-time, depending on extract | Household identity, rejected applications, technical eligibility, actual metered generation |
| Class county-year file | Beginner-ready deployment panel | projects, installed kWdc, estimated annual kWh | county-year | Building matching, peer exposure, household mechanisms |
| Google Solar faculty snapshot | Sampled technical opportunity | match quality, imagery date, roof area/segments, sunshine, modeled layouts and production | aggregate vetted records to ZIP/county | Installed panels, adoption date, household NPV, representative statewide rates |
| American Community Survey | Area conditions and denominator | housing units, owner occupancy, income, tenure, structure type, demographics | tract/ZIP/county with compatible years | Adopter characteristics or individual behavior |
| TIGER/Line | Geographic crosswalk and spatial structure | boundaries, GEOIDs, adjacency, land/water area | stable geographic identifiers | Economic mechanism or causal effect |
| Policy, tariff, and market records | Mechanism and comparison | incentive blocks, rates, compensation, program dates, installer presence, permits | utility/place-time | Credible causality unless timing and comparison assumptions are defensible |
The Lecture 3 audit reports 10,000 candidate points, 8,311 API returns, 4,829 valid matches within 30 meters, and 4,791 valid matches with potential fields. New York City supplies 67.45% of valid matches; 531 valid returned buildings lie outside New York; imagery dates span 2012–2024. Those facts imply four rules:
input_id values.Stock with stock, flow with flow. Compare cumulative installed capacity through a common date with a technical-capacity stock. Use annual project counts for policy timing or diffusion—not as the numerator of a cross-sectional “realized share.”
Question. How does realized distributed-solar deployment vary with sampled rooftop opportunity across New York places, and how do income, tenure, housing form, and regional market context alter that relationship?
Unit. Begin at county-year with the open class file. Move to ZIP-year if the full NYSERDA extract, stable geography, and adequate cell sizes support it. Aggregate Google buildings to the same geography; never infer a direct adopter match.
Primary outcomes. Keep separate:
Technical-opportunity measures. Report sampled valid buildings, potential capacity/output, roof/imagery coverage, and the share of candidate points with valid returns. The denominator must describe the sample; it is not the number of all eligible buildings in the geography.
Interpretation. The first result is a potential–deployment gradient or a map of deployment relative to sampled technical opportunity. It is not an inefficiency estimate. County and regional differences are descriptive until a design supplies an exogenous comparison.
| Hypothesis | Observable implication | Rival explanation | Evidence that would move the claim forward |
|---|---|---|---|
| H1 · Technical opportunity | Places with greater suitable sampled capacity have more cumulative deployment | Google coverage and urban sampling | A representative building frame, coverage weights, and sensitivity to imagery/match quality |
| H2 · Private value | Deployment differs across utilities or tariff regimes at comparable potential | Income, demand, housing, and local policy | Historical tariffs, net-metering rules, consumption proxies, and boundary/time variation |
| H3 · Liquidity / tenure | Lower deployment in low-income or renter-heavy areas at comparable opportunity | Building type, roof condition, or installer targeting | Financing eligibility/rollout, parcel tenure, and program take-up or application data |
| H4 · Peer learning | Prior nearby installations predict later adoption | Sorting, shared shocks, policy, and installers | Fine time/space histories plus an instrument, boundary, campaign, or other source of exogenous exposure |
| H5 · Installer access | Low deployment where few developers operate or quotes arrive | Low underlying demand | Quote/lead data, installer entry, travel costs, or market-boundary changes |
| H6 · Policy timing | Installations bunch or change around incentive blocks | Anticipation, seasonality, prices, and administrative delay | Exact rule dates, eligibility, application/completion timing, and untreated comparison places |
Where and when did projects, capacity, and sampled potential occur?
NYSERDA + vetted Google + transparent denominatorsHow do patterns differ with income, tenure, housing, region, or installer presence?
Add ACS, geography, coverage, and uncertaintyWhat changed adoption because of a policy, information intervention, or market-access shift?
Add treatment timing, a counterfactual, and design-specific assumptionsDid the change improve private or social welfare, and for whom?
Add prices, real costs, transfers, grid effects, emissions, and distributionMove upward only when the data support the next claim. A sophisticated model does not substitute for a credible comparison.
The first project cycle can remain entirely within open data while the restricted Google snapshot is reviewed.
Audit county-year rows, units, missingness, time coverage, and cumulative versus annual outcomes.
Use owner-occupied or suitable housing units when the claim requires an opportunity set; explain why it matches the outcome.
For each proposed variable, name the theory, predicted sign or comparison, and at least one rival mechanism.
Filter valid New York parents, quantify coverage and imagery vintage, and aggregate without exposing coordinates or identifiers.
Start with county or ZIP summaries; preserve a table of every crosswalk and unmatched unit.
Re-estimate under alternative years, denominators, coverage thresholds, regional exclusions, and outcome definitions.
Document rows, units, missingness, geography, time, duplicates, and stock/flow construction.
Trace one claim—tariffs, finance, peers, installer access, tenure, or equity—from theory to estimand and limitation.
Record incentive blocks, rate changes, Solarize campaigns, and program rules with dates and geographic eligibility.
Test how maps change with normalization, coverage, region, and the Modifiable Areal Unit Problem.
Translate a paper’s unit, outcome, sample restriction, and comparison into a public-data approximation.
Separate public tables from restricted coordinates, credentials, raw responses, and derivative exports.
For every result, record:
Begin with a transparent descriptive baseline. Add data only when they distinguish an economic explanation or unlock a more credible comparison.