Run Adjusted Linear Model¶
Goal¶
Fit adjusted linear models for one or more outcomes, then save coefficient tables, model summaries, adjusted group means, adjusted-mean contrasts, figures, a manifest, and an overview montage.
Use linear_model for new adjusted-model
workflows. Use adjusted_correlation
when the question is residualized correlations, not adjusted means.
Inputs¶
- A
Batch, experiment-like object,DataFrameExperiment, or raw DataFrame. - Outcome columns through
data_cols,outcomes, ordependent_variables. - A primary group column when adjusted means are wanted, such as
Diagnosis. - Predictor/covariate columns such as
Age,Sex, orsleep treatment. - Categorical/reference-level choices for interpretable coefficients.
Minimal Path¶
from PyFLASH import linear_model, load_state
batch = load_state(r"C:\path\to\pickles\SCN_Diagnosis.pkl")
result = linear_model(
batch,
data_cols=["Total counts"],
group="Diagnosis",
predictors=["Age", "Sex"],
categorical=["Sex"],
reference_levels={"Diagnosis": "Control", "Sex": "Female"},
save=False,
)
print(result["coefficients"].head())
print(result["adjusted_means_table"].head())
Full Workflow¶
- Confirm the outcome, group, and predictor columns exist:
needed = ["Total counts", "Amplitude", "Diagnosis", "Age", "Sex"]
print(batch.summary[needed].head())
- Start with a dry run. This catches formula, categorical, and missing-column issues before writing files.
dry = linear_model(
batch,
outcomes=["Total counts", "Amplitude"],
group="Diagnosis",
predictors=["Age", "Sex", "sleep treatment"],
categorical=["Sex"],
reference_levels={"Diagnosis": "Control", "Sex": "Female"},
interactions=[("Diagnosis", "Sex")],
save=False,
)
print(dry["model_summaries"])
print(dry["adjusted_mean_comparisons"])
- Choose adjusted-mean behavior. The default profile is
mean_mode; estimated marginal means usecovariate_profile="emm"or"reference_grid".
emm = linear_model(
batch,
data_cols=["Total counts"],
group="Diagnosis",
predictors=["Age", "Sex"],
categorical=["Sex"],
reference_levels={"Diagnosis": "Control"},
covariate_profile="emm",
adjusted_mean_weights="equal",
save=False,
)
- Save a reproducible run folder:
saved = linear_model(
batch,
outcomes=["Total counts", "Amplitude"],
group="Diagnosis",
predictors=["Age", "Sex", "sleep treatment"],
categorical=["Sex"],
reference_levels={"Diagnosis": "Control", "Sex": "Female"},
interactions=[("Diagnosis", "Sex")],
run_label="diagnosis_adjusted_models",
if_exists="version",
save=True,
)
print(saved["fig_dir"])
print(saved["adjusted_means_dir"])
- Run model-only tables when no group-adjusted means are needed:
model_only = linear_model(
batch,
data_cols=["Amplitude"],
predictors=["Age", "Sex"],
categorical=["Sex"],
adjusted_means=False,
plot_adjusted_means=False,
plot_coefficients=False,
save=False,
)
- Use
filter_byto model a row subset or queue. Pipeline queues share one run folder and tag each filter's files.
female_male = linear_model(
batch,
data_cols=["Total counts"],
group="Diagnosis",
predictors=["Age"],
filter_by=[
{"Sex": "Female"},
{"Sex": "Male"},
],
run_label="diagnosis_by_sex_models",
)
Outputs¶
With save=True, the run folder is:
Common outputs include:
| Output | Meaning |
|---|---|
linear_model_coefficients.csv |
Estimates, intervals, p-values, and corrected q-values. |
linear_model_summaries.csv |
Model-level fit summaries and counts. |
linear_model_metadata.csv |
Resolved formulas, terms, and fit status. |
Adjusted Means/linear_model_adjusted_means.csv |
Model-adjusted group means. |
Adjusted Means/linear_model_adjusted_mean_comparisons.csv |
Pairwise adjusted-mean contrasts. |
Coefficient Forest.svg |
Coefficient summary figure. |
Adjusted Means/*.svg |
Adjusted-mean figures by outcome. |
manifest.json |
Stable run summary. |
../_runs_index.csv |
One-row-per-run index for linear-model runs. |
! Overview Montage.png |
Overview montage when saving and montage capture is enabled. |
The Python result dictionary also includes the same tables as in-memory DataFrames for fresh runs.
Troubleshooting¶
linear_model adjusted means need group=<column>: pass a group column, or set bothadjusted_means=Falseandplot_adjusted_means=False.- Categorical levels look wrong: pass
categorical=[...]andreference_levels={...}explicitly. - A predictor is missing because of display-name changes: inspect
batch.summary.columns; PyFLASH resolves some aliases but not arbitrary Excel labels. - Too many coefficient terms in the forest plot: use
max_coefficient_termsor narrow the predictor set. - Need legacy table-only outputs under
Modelling/Linear Models: userun_linear_model_pipeline.