DATA SCIENCE
EDA: Austin Water Quality Samples
This analysis explored the City of Austin’s public Water Quality Sampling dataset, asking: how does depth and location affect the filter result of an Austin water sample? The raw dataset contained 303,730 rows and 24 variables spanning multiple decades of surface-water testing across the city.
Data cleaning: A large share of the raw data was unusable for this question. 206,950 rows were missing depth values and 145,978 were missing a sample ID, both of which were filtered out (rows lacking a sample ID couldn’t be verified). After cleaning and reshaping, the working dataset was reduced to 46,920 tidy observations across 135 distinct sampling sites.
Findings:
- Sample depth was heavily right-skewed, with the deepest sampling sites concentrated at Austin’s lakes (Lake Long, Lake Austin, Lady Bird Lake).
- Fecal coliform bacteria contamination varied widely by site (mean 204 colonies/100mL, max 1,624), with the highest levels found at Tributary 6 @ Bull Creek, Taylor Slough South @ Reed Park, and Lady Bird Lake @ 1st St.
- Ammonia contamination showed a clear depth relationship: concentrations dropped off sharply below 2.5 meters, with the highest average levels clustered at Lady Bird Lake and the Buda sites, including one site downstream of a known chemical spill.
- Both contaminants were more common, and more variable, at shallow depths, partially confirming the initial hypothesis that surface-adjacent samples would show different contamination patterns, though the data was too sparse at greater depths to draw a fully confident trend.
Takeaway: specific Austin locations, particularly certain lakes and tributaries, show meaningfully elevated contamination levels, pointing to a need for targeted water-quality monitoring in those areas rather than a citywide default assumption.
Completed for Elements of Data Science (SDS322E), under Prof. Layla Guyot and TA Walter Wang. Dataset via the City of Austin Open Data Portal (owner: Robert Clayton). Stack: R, tidyverse (dplyr, ggplot2, lubridate)
Stack: R, tidyverse (dplyr, ggplot2, lubridate)