Boston College · B.A. Economics, May 2027

KELLEN
KUBEC

I work on applied econometrics — building county-level panels out of public data, estimating causal effects, and testing whether the effect survives a better specification. Right now I’m also shipping production software as a development intern.

Kellen Kubec
3.75 / 4.00 · Dean’s ListB.A. Economics, minors in Mathematics and Finance
Two econometrics papersDifference-in-differences and maximum-likelihood work, both on original data
Stata, R, SQL, Next.jsThree years in Stata; full-stack development on a live client product

Selected research

Two papers, two questions about whether an estimate means what it appears to mean.

Difference-in-differences · county–month panel, 2012–2017

Impact of Hurricane Matthew on Employment Across Income Groups of the Carolinas

Hurricane Matthew raised unemployment by 0.475 percentage points in the three months after landfall — but the average hides the result that matters. High-income counties absorbed no measurable employment shock at all, and that gap was still there fifteen months later, after the average effect had disappeared.

+0.475pp average effect on unemployment, 3-month window (p = 0.045)
+0.618pp for middle-income counties, the reference group
−0.797pp top-tercile differential — wealthy counties insulated (p < 0.05)
145counties, 8,700 county-months in the 3-month sample
Abstract

This paper examines how the increase in natural disaster frequency has intensified the conversation about economic disruptions and raised questions about unequal labor market effects across communities, focusing on the impact of Hurricane Matthew in North Carolina and South Carolina. Using monthly county unemployment data, median income measures, and NOAA storm intensity information, we apply a difference in differences framework that compares affected counties to unaffected counties before and after landfall. We estimate both an overall hurricane employment effect and variation across income terciles while also assessing short run and longer horizon responses. Our results show that Hurricane Matthew increased unemployment by about 0.48 percentage points on average and by 0.63 percentage points for middle income counties, while low-income counties show no significant difference from middle income counties, high-income counties show to be insulated from unemployment effects and experience no significant change in unemployment, suggesting faster adjustment and recovery. These findings deepen evidence on the distribution of climate driven labor shocks and indicate that employment displacement falls unevenly across income groups, informing the design of targeted post disaster support for workers in areas with weaker recovery capacity.

Implied treatment effect by income tercile

Three months after landfall. Point estimates with ±1.96 standard errors, clustered by county (Table 2, Column 6).

−1 0 +1 +2 Change in unemployment rate (percentage points) Bottom 33% +1.06 Middle 33% +0.62** Top 33% −0.18

Only the middle tercile’s implied effect is individually significant (** p < 0.05). The bottom tercile’s point estimate is large but imprecise. What is significant is the gap: the Treated × Post × Top33 interaction is −0.797 (p = 0.016), and it barely moves at the 15-month horizon (−0.753, p = 0.032).

Difference-in-differences County / year / month fixed effects SEs clustered by FIPS Parallel-trends diagnostics Seasonal adjustment BLS LAUS U.S. Census SC Revenue & Fiscal Affairs NOAA windspeed Stata
Figures, diagnostics and regression tables
GroupnMean unemp.Avg. median household income
Treated1,8007.78%$44,502.68
Untreated8,6407.23%$44,460.59
Income group 1 (low)3,5508.68%$35,895.45
Income group 2 (middle)3,5066.97%$43,543.63
Income group 3 (high)3,3846.27%$54,418.29
Treated × income 17209.13%$35,202.50
Treated × income 24216.93%$44,937.79
Treated × income 36596.85%$54,385.75
Untreated × income 12,8308.65%$36,071.74
Untreated × income 23,0856.98%$43,353.38
Untreated × income 32,7256.13%$54,426.16
Figure 1. Summary statistics by income tercile and treatment status, dataset including 2017. All dollar figures converted to 2025 USD. Income terciles are assigned on 2016 median household income so counties cannot switch groups over the panel.
Regression equation used to estimate month-of-year seasonal effects
Figure 2. The seasonal-adjustment regression. Because coastal Carolina counties hire heavily in the summer, unemployment was residualized on month-of-year within each income × treatment cell before the parallel-trends check, then the cell mean added back.
Parallel trends plot for low-income counties
Figure 3. Parallel trends, low-income counties. Treated in red, untreated in blue; the dotted vertical line is September 2016.
Parallel trends plot for middle-income counties
Figure 4. Parallel trends, middle-income counties.
Parallel trends plot for high-income counties
Figure 5. Parallel trends, high-income counties.
Parallel trends plot comparing all six income-by-treatment groups
Figure 6. All six income × treatment cells on one axis. The differences between groups stay roughly constant through September 2016, which is what the design needs.
The six difference-in-differences specifications estimated in the paper
Figure 8. The six specifications. Equations 1–3 estimate an average effect, adding clustering and then county, year and month fixed effects; equations 4–6 repeat the ladder with income-group interactions. Equations 3 and 6 are the preferred models.
Regression table: three-month post-hurricane window
Table 2. Three-month window — the main results. Panel E reports the implied effect for each tercile.
Regression table: fifteen-month post-hurricane window
Table 1. Fifteen-month window. The average effect falls to 0.249 and loses significance (p = 0.253) — but the top-tercile differential holds at −0.753 (p = 0.032). Opposing effects across income groups cancel in the aggregate.

Heteroskedastic & skew probit · 95,781 NCAA games, 2003–2025

Spread Size, Outcome Variance, and the Illusion of Point Shaving in NCAA Basketball

Wolfers (2006) reads the gap between two spread-coverage probabilities as evidence that college basketball players shave points. The gap is arithmetic: it appears just as clearly in simulated games drawn from a symmetric normal distribution where, by construction, no manipulation exists.

95,781Division I games, 2003–04 through 2024–25
42forecasting models averaged into the market spread
−70%drop in the spread coefficient once outcome variance is modeled
70,207games in the estimation sample (favorite won)
Abstract

This paper examines whether the divergence documented by Wolfers (2006) can identify point shaving in NCAA basketball. Wolfers uses the gap between a favored team’s probability of winning without covering the point spread and its probability of winning by twice the spread as evidence of systematic manipulation in collegiate athletics. Using 95,781 Division I games spanning 2003 to 2025, I show that this identification strategy is fundamentally flawed, because the divergence between these two probabilities is a mechanical consequence of spread size rather than evidence of manipulation. Specifically, a team favored by S points has exactly S−1 winning margins that fail to cover, so the probability of winning without covering rises mechanically with spread size even in the complete absence of point shaving. A simulation exercise confirms this formally: when margins of victory are drawn from a normal distribution centered on the spread with no manipulation imposed, the Wolfers divergence appears clearly in the simulated data.

Correcting for this asymmetric mechanical relationship with a heteroskedastic probit that explicitly models the increasing variance of game outcomes in spread size, I find no systematic evidence that larger favorites are more likely to win without covering, and estimating the model year by year across the two-decade sample reveals no time trend in the residual effect of spread size.

What happens to the “point shaving signal” when variance is modeled

Coefficient on the standardized absolute spread, in the mean equation. Dependent variable = 1 if the favorite won but failed to cover.

0 .25 .50 .75 Coefficient on z_absline Standard probit 0.668 t = 87.09 Heteroskedastic probit 0.197 t = 9.11

The variance equation does the work: spread size predicts outcome dispersion with a coefficient of 1.009 (t = 23.45). Blowouts are noisier — garbage time, pulled starters, slackening defense — and once that is absorbed, most of the apparent signal goes with it.

Maximum likelihood Heteroskedastic probit Skew probit (Azzalini 1985) Monte Carlo simulation Kernel density estimation Model selection by log-likelihood Stata
Figures and model comparison
Summary statistics by season
Table 1. Summary statistics by season. Average spread holds between 7.3 and 8.2 points across two decades; the share of games where the favorite won without covering drifts from roughly 24–26% down to 21–23%.
Kernel density of winning margin relative to the spread, games with spreads under twelve
Figure 1. Kernel density of the forecast error — realised margin minus spread — for games with spreads under twelve. Close to normal, as a well-priced market should be.
Kernel density of winning margin relative to the spread, heavy favorites
Figure 2. The same density for heavy favorites, which is slightly right-skewed — the pattern Wolfers reads as “too few” big favorites covering.
Wolfers-style divergence of the two probability series, binned by spread
Figure 3. The Wolfers comparison replicated on 2003–2025 data: P(won but did not cover | won) against P(did not beat twice the spread | covered), binned by spread. The divergence is clearly present.
Wolfers-style divergence estimated separately across five time periods
Figure 4. The same divergence across five periods. The pattern repeats throughout, though it narrows after 2019 — which a Wolfers reading would call declining manipulation.
The same divergence reproduced in fifty simulated samples containing no manipulation
Figure 5. The central result. Fifty simulated replications of the full sample, with margins drawn from a normal distribution centered on the spread and zero manipulation by construction. The divergence appears anyway — so it cannot be evidence of point shaving.
Model comparison: probit, heteroskedastic probit and two skew-probit specifications
Table 2. Four specifications on the 70,207-game sample. The spread coefficient falls from 0.668 to 0.197 once variance is modeled. The flexible skew probit fits best (log-likelihood −38,962.4 against −39,438.4 for the standard probit), which says the symmetry assumption underlying the standard model is itself rejected by the data.

Coursework project · Stata and Excel

Maximum likelihood from first principles, and a Bradley–Terry rating system

Estimated Logit and Probit models by maximum likelihood to predict NCAA game outcomes, computed marginal effects, and benchmarked them against OLS and truncated linear-probability specifications. Built the log-likelihood functions by hand in Excel and optimized the parameters with Solver — useful mostly because it makes it obvious what a canned logit command is actually doing.

Separately, constructed a 32-team Bradley–Terry paired-comparison rating system for NFL performance and estimated it under both OLS and Logit.

Why this is here

Both papers above lean on likelihood-based estimation and on knowing when a default specification is quietly assuming something false. This is where that came from.

Experience

Software in production, and budgets that had to be right.

June 2026 – present
Remote

Software Development Intern

Integryl

  • Developer behind a production web application used daily by a general contracting client, translating operational workflows into shipped software.
  • Built full-stack functionality on Next.js and Supabase/PostgreSQL, including role-based access control and a client-facing portal.
  • Ran the full build cycle — scoping, development, debugging, code review, deployment — in an AI-native workflow, validating every output before release.
  • Gathered requirements directly from the client, triaged issues, and iterated against live operational feedback.
Summer 2025
Hilton Head, SC

Budgeting Analyst Intern

RCH Construction, Inc.

  • Prepared budgets for residential construction projects, estimating material and labor requirements from design parameters.
  • Forecast project costs against prevailing and projected market conditions for materials and labor.
  • Generated and analyzed variance reports comparing actual spend to budget, isolating the drivers of cost deviation.
Summer 2024
Hilton Head, SC

Project Manager Intern

RCH Construction, Inc.

  • Supported project managers across construction phases; applied quality control and efficiency metrics to active builds.
  • Coordinated and scheduled subcontractor work across multiple concurrent projects.

Background

Education, toolkit and the rest.

Education
Boston College — Chestnut Hill, MAB.A. Economics; minors in Mathematics and Finance. Expected May 2027.
3.75 / 4.00 cumulative · Dean’s ListCoursework: Applied Econometrics, Honors Microeconomics, Honors Macroeconomics, Investments and Derivatives, Money and Financial Markets.
Toolkit
EconometricsStata (3+ years), R (2+ years), SQL. Advanced Excel, including maximum likelihood via Solver.
SoftwareNext.js, PostgreSQL / Supabase, AI-assisted development with Claude Code. Adobe Illustrator.
Awards
RBC Heritage Scholarship2023–2027
Krum Foundation Scholarship2023–2027
Mu Alpha ThetaNational Mathematics Honor Society · AP Scholar
Leadership & interests
Vice President, Boston College Chess Club2023 – present
Team Captain, Hilton Head Rowing Club2021 – 2023
OtherwiseChess, guitar, rowing.