Project Documentation¶
This site provides project documentation. Use the documentation navigation to explore.
How-To Guide¶
Many instructions are common to all our projects.
See ⭐ Workflow: Apply Example to get the example projects running on your machine.
Project Documentation Pages (docs/)¶
- Home - this documentation landing page
- Project Instructions - the standard project workflow
- Your Files - how to copy the example and create your version
- Glossary - project terms and concepts
- API - autogenerated code documentation for the public project interface
Serving a Model (Hosting an Endpoint)¶
Deployment Options¶
- HuggingFace - free, no CC required
- Render - free, easier, CC required
Phase 4. Technical Modification¶
I copied serve_case.py to serve_gracecode42_modification.py and changed what the serving core returns. The original sends back only the predicted species. Mine also returns the probability for each of the three classes, from model.predict_proba. Another update to the same function replaced the plain list passed to model.predict with a one-row DataFrame carrying the four feature names, and I renamed the logger to M06-MOD so my runs are identifiable in project.log.
The example notebook offered this as a direction to explore. It was a simple change to make and the added information is interesting.
I verified the change by running my file in a terminal, sending the three payloads I had used during setup, and rerunning my copy of the notebook top to bottom. The three requests returned probabilities of 0.995, 1.0, and 1.0 for their predicted species. The notebook's payload returned Chinstrap at 0.83, with 0.145 for Adelie and 0.025 for Gentoo. In the original version all four payloads return a species, so nothing separated the payload at 0.83 from the ones at 1.0. The sklearn warning about missing feature names, which printed on every request before, is gone from both the server log and the notebook output.
The code change was about six lines. Most of the work was updating docstrings, run commands, and the logger tag across the two files, so I would call it easy to moderate. Ruff rejected my first version for calling zip without saying what should happen if the two sequences differed in length, and strict=True resolved it.
Phase 5. Custom Project¶
Basis and Data¶
The example serves a penguin species classifier trained on four body measurements. I replaced the dataset with medical insurance charges, 1,338 individuals with age, sex, BMI, number of children, smoking status, region, and annual charges in dollars, from the Medical Cost Personal Datasets on Kaggle under ODbL v1.0.
I chose it because I had already worked with it twice. Module 4 fit regression models to charges and Module 5 built a classifier on cost tiers, and both projects ended with questions I did not have time to answer. This module gave me the chance to finish that work and serve the result.
The provenance is not fully documented, so the dataset is best treated as representative rather than authoritative. Its more important limitation is that it records no diagnosis, procedure, or claims history, which turns out to determine what the models can and cannot do.
Example Model and Serving Approach¶
The example trains a RandomForestClassifier on the penguins dataset in model_builder_case.py, saves it to artifacts/model.joblib with joblib, and serves it from serve_case.py. FastAPI loads the artifact once at startup, and a POST to /predict carrying four measurements returns the predicted species. My Phase 4 modification added the probability for each class to that response.
Loading at startup rather than per request is why the builder and the server are separate files. The builder's only product is the saved artifact, and the server never trains anything.
Custom Application¶
I serve two models from one request. A POST to /predict carries four fields, age, bmi, children, and smoker, and the response returns a cost tier with the probability of each of the four tiers alongside a predicted dollar amount. The tiers come from cutting charges at its quartiles, at 4,740, 9,382, and 16,640 dollars.
Both models are pipelines rather than bare estimators, so the feature derivation travels inside the saved artifact and the server cannot compute a column differently than training did. The derivation lives in features_gracecode42_project.py, which the builder, the server, and the notebook all import, so there is one definition rather than three.
The classifier is a random forest with 200 trees at depth 10, using age, bmi, children, smoker, and an engineered smoker_bmi term. The regressor is linear regression on those same fields plus smoker_bmi and smoker_bmi30, a flag for smokers at BMI 30 or above.
The two engineered terms are not interchangeable, and which one helps depends on the model. A tree splits one feature at a time and cannot build a product, so smoker_bmi gives it something new; smoker_bmi30 gains it nothing, because finding thresholds is what a tree already does. A linear model is the reverse. It builds products freely once expanded but cannot produce a step at all, so smoker_bmi30 is what it needs. Testing features and models as a grid rather than in sequence is what made that visible.
Every choice was made by five-fold cross-validation on the 1,070 training rows. The 268 test rows were held back until both models were fixed and scored once.
I verified the work three ways. The notebook calls the same prediction function the server calls, in-process, on four payloads and on three deliberately bad ones, which return HTTP 400 from the server and a plain ValueError in the notebook. The builder reports its test scores in project.log. And I sent the payloads to the running endpoint with curl and compared the responses.
The model is deployed locally, started with
uv run fastapi dev src/mlstudio/serve_gracecode42_project.py.
Summary¶
The classifier reaches 0.8582 accuracy and 0.8556 weighted F1 on the held-out 268. The regressor reaches R-squared 0.8609 with RMSE 4,607.9 dollars, close to the 4,551 I reached in Module 4 while using six terms instead of forty-four.
The clearest result is that reasoning about features beat adding model complexity. Plain linear regression on four fields gives cross-validated RMSE of 6,049 dollars. Expanding those fields to degree 2 builds fourteen terms and reaches 4,863. Adding smoker_bmi as one column reaches 4,873 with five terms, matching the polynomial, because a degree-2 expansion already contains that product. Adding smoker_bmi30 reaches 4,473, and both together reach 4,412 with six terms. Six terms beat twenty-seven.
The residual plot is what pointed to smoker_bmi30. Plain linear regression split its errors into two flat bands, smokers under BMI 30 at a median of negative 9,716 dollars and smokers at BMI 30 or above at positive 7,587, because one smoker coefficient has to cover both groups. Flat bands mean the premium jumps at a threshold rather than rising with BMI, and a step is what a polynomial cannot represent no matter how many terms it is given.
What neither model can do is the more useful finding. The very high tier holds near 0.76 recall in every cell detailed in Section 7b and falls to 0.7164 on the test set, and nothing among the features or models tried moves it. Eight of the 268 test cases sit in the top tier and were predicted low, three tiers down. Twenty-three cases have residuals above 7,000 dollars, averaging 14,218, and twenty of those are non-smokers. These overlap heavily: the classifier's misses and the regressor's largest under-predictions are drawn from the same eighteen non-smokers sitting in the top tier. Age, BMI, children, and smoking do not distinguish them from people who cost a fraction as much, so no term built from those columns will reach them. That is a limit of the data rather than a failure of the model, and saying so is more honest than adding complexity that cannot help.
I would apply this to student outcomes at the college where I work, where the same shape appears: a tier is what a person acts on, a number is what gets budgeted, and the cases a model misses badly are usually the ones whose cause was never recorded.
