Daily energy against average temperature
Correlation is a question about two measures, so both go on axes and neither goes on time - the days become the points.
How it decided
Energy readings: 6,483 rows, 4 columns. temperature_c and kwh per day, split by site.
6,483 rows in, 273 out across 4 columns. Correlation is a question about two measures, so both go on axes and neither goes on time - the days become the points.
scatter. Two measures and no time axis is a scatter. Three sites is the cap for a scatter here, because adjacent points must stay separable for a colour-blind reader.
Palette, spacing, axis titles and legend placement applied. 0 issues found.
The transformation
This is the code, not a description of it — the chart above was drawn from what it returns, and the notebook download runs the same lines against the same file.
df["day"] = df["reading_at"].astype("datetime64[ns]").dt.floor("D")
result = df.groupby(["day", "site"], as_index=False).agg(
kwh=("kwh", "sum"), temperature_c=("temperature_c", "mean")
)
What it charted
273 rows out of 6,483, 4 columns. First 8 shown.
| day | site | kwh | temperature_c |
|---|---|---|---|
| 2025-01-01 | Plant A | 24167.81 | 10.1417 |
| 2025-01-01 | Plant B | 13349.33 | 10.075 |
| 2025-01-01 | Plant C | 36078.12 | 11.8583 |
| 2025-01-02 | Plant A | 24128.19 | 11.025 |
| 2025-01-02 | Plant B | 13627.38 | 10.8667 |
| 2025-01-02 | Plant C | 36976.89 | 10.0083 |
| 2025-01-03 | Plant A | 24382.17 | 10.9208 |
| 2025-01-03 | Plant B | 13572.79 | 11.0917 |