twoHelixestwoHelixes Start free

Datasets / scikit-learn

Diabetes progression

Ten baseline variables against disease progression after a year.

442 rows11 columns scikit-learn

Worked examples

Each of these is a real run: the pipeline picked the columns, shaped the rows and chose the form. Open the trace to see what it decided and why.

Baseline variables ranked by correlation with progression

Absolute correlation, so a strong negative relationship is not sorted to the bottom as though it were weak.

Correlation with disease progression00.20.40.6bmis5bps4s3s6s1ages2sexCorrelation
which variable correlates most with progression? — 10 rows charted from 442, drawn as a hbar
How it decided
Finding the data139 ms

Diabetes progression: 442 rows, 11 columns. all ten baseline variables against progression.

Shaping the data14 ms

442 rows in, 10 out across 2 columns. Absolute correlation, so a strong negative relationship is not sorted to the bottom as though it were weak.

Choosing the chart9 ms

hbar. A ranking across ten items: horizontal bars, strongest first.

Applying defaults1 ms

Palette, spacing, axis titles and legend placement applied. 0 issues found.

BMI against disease progression

One point per patient, no aggregation: the spread around the trend is as much of the answer as the trend.

BMI against progression01002003000.061696206518683…-0.02991781976118…-0.07949717515970…Bmi
show BMI against progression — 442 rows charted from 442, drawn as a scatter
How it decided
Finding the data139 ms

Diabetes progression: 442 rows, 11 columns. bmi and progression.

Shaping the data7 ms

442 rows in, 442 out across 2 columns. One point per patient, no aggregation: the spread around the trend is as much of the answer as the trend.

Choosing the chart2 ms

scatter. Two measures, no category and no time: a scatter of the raw rows.

Applying defaults1 ms

Palette, spacing, axis titles and legend placement applied. 0 issues found.

Schema

11 columns, typed as the pipeline sees them — which is how it knows what can go on a time axis and what can be summed.

Column TypeExample
agenumber0.0381
sexnumber0.0507
bminumber0.0617
bpnumber0.0219
s1number-0.0442
s2number-0.0348
s3number-0.0434
s4number-0.0026
s5number0.0199
s6number-0.0176
progressionnumber151.0

First 8 rows

agesexbmibps1s2s3s4s5s6progression
0.03810.05070.06170.0219-0.0442-0.0348-0.0434-0.00260.0199-0.0176151.0
-0.0019-0.0446-0.0515-0.0263-0.0084-0.01920.0744-0.0395-0.0683-0.092275.0
0.08530.05070.0445-0.0057-0.0456-0.0342-0.0324-0.00260.0029-0.0259141.0
-0.0891-0.0446-0.0116-0.03670.01220.025-0.0360.03430.0227-0.0094206.0
0.0054-0.0446-0.03640.02190.00390.01560.0081-0.0026-0.032-0.0466135.0
-0.0927-0.0446-0.0407-0.0194-0.069-0.07930.0413-0.0764-0.0412-0.096397.0
-0.04550.0507-0.0472-0.016-0.0401-0.02480.0008-0.0395-0.0629-0.0384138.0
0.06350.0507-0.00190.06660.09060.10890.02290.0177-0.03580.003163.0