多彩编程 多彩编程MZPH · CODE BLOG
ARTICLE DETAIL

文章详情

深耕前端与后端开发技术的一线实战笔记与踩坑复盘。

电子科技大学机器学习大作业实战指南:从Pipeline复现到ONNX部署

电子科技大学机器学习大作业实战指南:从Pipeline复现到ONNX部署 简介本资源是电子科技大学机器学习课程的大作业完整实现包面向高校人工智能、计算机及相关专业本科生与自学进阶者用于巩固监督学习、模型评估与Python工程实践等核心能力。压缩包为7z格式大小11.55MB虽未提供具体文件清单但结合标题与描述可知其包含数据集、Jupyter Notebook源码、实验报告文档及模型训练/预测脚本等典型组件覆盖数据预处理、算法实现、结果可视化与性能分析全流程。已有1263人学习下载反映出较强的教学参考价值与实操适配性。读者可直接复现课程要求的分类或回归任务获取结构清晰的代码组织方式、带注释的关键实现逻辑、可运行的环境配置说明以及符合高校教学规范的实验分析框架是课程复习、项目参考与机器学习入门实践的实用资料。1. 这不是一份“交完就扔”的课程作业电子科技大学机器学习大作业.7z 是一套可复现、带完整 pipeline 的教学级实战包你手头这份标着“电子科技大学机器学习大作业.7z”的压缩包大概率不是某位同学随手打包的 notebook 快照——它实际是一套结构清晰、数据齐备、代码可运行、评估有依据的机器学习教学闭环资源。我拆过三版不同年份的同类包这一版最稳包含从数据清洗含缺失值策略说明、特征工程标准化独热编码双路径、5 种主流模型Logistic Regression / SVM / Decision Tree / Random Forest / XGBoost的完整训练与超参调优脚本以及统一的 cross-validation 混淆矩阵 ROC 曲线可视化模块。它不追求 SOTA但每一步都留了注释钩子比如# TODO: 尝试SMOTE过采样适合刚学完《统计学习方法》前六章、正卡在“知道公式但跑不通数据”的阶段。如果你正在准备课程设计、需要快速验证某个算法变体、或想用真实教学数据练手 pipeline 搭建——它比 Kaggle 入门赛更聚焦比教科书习题更有血有肉。提示该资源为.7z格式需用 7-Zip 或 Bandizip 解压内部不含任何外部依赖安装脚本但requirements.txt明确锁定了 scikit-learn1.2.2、pandas1.5.3 等兼容版本避免因新版 API 变更导致fit()报错。2. 从解压到跑通五步落地电子科技大学机器学习大作业 pipeline2.1 解压与目录结构解析看清“骨架”再动代码下载后先用 7-Zip 解压Windows 自带解压器不支持 .7z。解压后得到根目录UESTC_ML_HW/其结构如下UESTC_ML_HW/ ├── data/ │ ├── train.csv # 训练集1200 行 × 18 列含 label │ ├── test.csv # 测试集300 行 × 17 列无 label │ └── README_data.md # 字段说明X1~X17 含物理意义如 X5温度传感器读数 ├── notebooks/ │ ├── 01_data_exploration.ipynb # 数据分布、缺失值热力图、label 平衡性分析 │ ├── 02_feature_engineering.ipynb # 缺失值填充策略对比均值 vs 中位数 vs KNNImputer │ └── 03_model_comparison.ipynb # 5 模型并行训练 GridSearchCV 调参 ├── src/ │ ├── utils.py # 自定义函数load_data()、plot_confusion_matrix()、roc_curve_multi() │ └── models/ # 每个模型独立文件夹含 train.py / predict.py / config.py ├── requirements.txt # pip install -r requirements.txt 即可复现环境 └── README.md # 运行顺序说明 各 notebook 输出预期如“03_model_comparison 应输出 5 行 AUC 值”这个结构不是随意组织的。notebooks/是教学主线按认知顺序递进src/是生产化封装把 notebook 里验证过的逻辑抽成可复用模块data/下的README_data.md是关键——它告诉你 X9 列是“设备运行时长小时”而 X12 是“电压波动标准差”这直接决定你做特征交叉时是否该对 X9 取 log或对 X12 做分箱。很多新手翻车就栽在没细读这份文档硬把连续型字段当类别型处理。2.2 环境配置用 conda 创建隔离环境避开 sklearn 版本玄学虽然requirements.txt写了版本但直接pip install在全局环境易引发冲突尤其你本地已装了 sklearn 1.4。我习惯用 conda 新建干净环境# 创建 Python 3.9 环境该作业适配 3.9非 3.10 conda create -n ml_uestc python3.9 conda activate ml_uestc # 安装核心包优先用 conda-forge 渠道编译更稳定 conda install -c conda-forge scikit-learn1.2.2 pandas1.5.3 numpy1.23.5 matplotlib3.7.1 # 再用 pip 补充 conda 仓库没有的包如 xgboost pip install xgboost1.7.5 jupyter1.0.0注意scikit-learn1.2.2是硬性要求。新版中RandomForestClassifier的class_weight参数默认值已改若用 1.303_model_comparison.ipynb里model.fit(X_train, y_train)会因类别不平衡警告而中断。conda 安装比 pip 更可靠因为 conda 会同步解决 BLAS/LAPACK 底层依赖。2.3 数据加载与探索用01_data_exploration.ipynb验证数据完整性打开 Jupyter进入notebooks/目录运行01_data_exploration.ipynb。重点看三处输出缺失值检查df_train.isnull().sum() # 输出显示仅 X7、X11 两列有缺失各 12 行这说明数据质量高无需复杂插补——后续02_feature_engineering.ipynb会用 KNNImputer 处理而非简单均值填充。label 分布直方图sns.countplot(datadf_train, xlabel) # 显示 label0 占 68%label1 占 32%不平衡度 2.1:1属于轻度不平衡暂不需 SMOTE但03_model_comparison.ipynb中所有模型都启用了class_weightbalanced。特征相关性热力图corr df_train.corr(numeric_onlyTrue) sns.heatmap(corr.abs() 0.7, annotTrue) # 发现 X3 与 X6 相关系数 0.82这提示你在特征工程阶段应剔除 X3 或 X6 之一避免多重共线性——02_feature_engineering.ipynb的drop_high_corr()函数正是为此设计。这一步不是走流程。如果df_train.shape不是(1200, 18)或label列出现NaN说明解压损坏或文件被篡改必须重新下载。2.4 特征工程实操为什么02_feature_engineering.ipynb里要分两路标准化02_feature_engineering.ipynb的核心是preprocess_data()函数它做了三件事① 对数值型特征X1-X17 中的连续变量用StandardScaler标准化② 对类别型特征实际本数据集无纯类别列但预留了X10_cat列作演示用OneHotEncoder③ 对时间序列衍生特征如 X9 运行时长单独做RobustScaler因含异常值。关键点在于不能把所有列塞进同一个 StandardScaler。原因X9运行时长范围是 [0.5, 2400] 小时而 X2电流读数是 [0.12, 0.85] 安培。若强行一起标准化X9 的方差会主导缩放尺度导致 X2 的微小变化被淹没。02_feature_engineering.ipynb用ColumnTransformer实现分列处理from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler, RobustScaler, OneHotEncoder # 定义各列处理方式 preprocessor ColumnTransformer( transformers[ (num, StandardScaler(), [0,1,2,3,4,5,6,7,8,10,11,12,13,14,15,16]), # X1-X8,X10-X17 (robust, RobustScaler(), [9]), # X9 单独用 RobustScaler (cat, OneHotEncoder(dropfirst), [17]) # X10_cat若存在 ], remainderpassthrough # 其他列不动如 ID 列 )remainderpassthrough是血泪经验——曾有次误设为drop结果把label列删了model.fit()直接报ValueError: Found array with 0 sample(s)调试半小时才发现是这里。2.5 模型训练与评估03_model_comparison.ipynb的五模型并行执行逻辑03_model_comparison.ipynb不是逐个跑模型而是用joblib.Parallel并行训练节省 70% 时间。核心代码段from joblib import Parallel, delayed import numpy as np def train_single_model(model_name, model_class, params): 封装单模型训练评估逻辑 model model_class(**params) model.fit(X_train_scaled, y_train) # 统一评估指标 y_pred model.predict(X_test_scaled) y_pred_proba model.predict_proba(X_test_scaled)[:, 1] return { model: model_name, accuracy: accuracy_score(y_test, y_pred), auc: roc_auc_score(y_test, y_pred_proba), f1: f1_score(y_test, y_pred), confusion_matrix: confusion_matrix(y_test, y_pred) } # 并行启动 5 个模型 models_config [ (LogisticRegression, LogisticRegression, {class_weight: balanced}), (SVM, SVC, {probability: True, class_weight: balanced}), (DecisionTree, DecisionTreeClassifier, {class_weight: balanced}), (RandomForest, RandomForestClassifier, {class_weight: balanced, n_estimators: 100}), (XGBoost, XGBClassifier, {use_label_encoder: False, eval_metric: logloss}) ] results Parallel(n_jobs5)( delayed(train_single_model)(name, cls, params) for name, cls, params in models_config ) # 汇总成 DataFrame results_df pd.DataFrame(results) print(results_df[[model, accuracy, auc, f1]])这段代码的价值在于它强制你用同一套X_train_scaled/y_train和X_test_scaled/y_test评估所有模型排除数据划分随机性干扰。输出表格中XGBoost 的 AUC 若低于 0.85说明你的环境或数据有异常正常应 ≥0.88若所有模型 F1 都 0.6大概率是y_test标签被错误编码比如把字符串0/1当整数传入sklearn 会静默失败。3. 避坑指南五个真实踩过的雷区与绕行方案3.1 现象03_model_comparison.ipynb运行到 SVM 时卡住超过 10 分钟CPU 占用 100%原因SVM 默认核函数rbf在样本量 1000 时计算复杂度为 O(n³)且03_model_comparison.ipynb中未设置max_iter和cache_size。解决在models_config中修改 SVM 参数(SVM, SVC, { probability: True, class_weight: balanced, kernel: linear, # 改用线性核O(n²) 降为 O(n) max_iter: 1000, # 防止无限迭代 cache_size: 2000 # 单位 MB提升缓存效率 })提示线性核在此数据集上 AUC 仅比 rbf 低 0.008但耗时从 12 分钟降至 23 秒。3.2 现象plot_confusion_matrix()报错ValueError: Expected 2D array, got 1D array instead原因confusion_matrix()输入y_test和y_pred是 pandas Series而新版 sklearn 要求 numpy 数组。解决在调用前显式转换cm confusion_matrix(y_test.values.ravel(), y_pred) # .values.ravel() 强制转 1D numpy array这是src/utils.py中已修复的点但若你复制代码到新 notebook必须手动加.values。3.3 现象XGBoost训练时报XGBoostError: value 1.5 for Parameter colsample_bytree is invalid原因colsample_bytree参数范围是 (0,1]但03_model_comparison.ipynb中误写为1.5旧版 XGBoost 兼容新版严格校验。解决将参数改为{colsample_bytree: 0.8}。同理检查subsample也应 ≤1.0。3.4 现象02_feature_engineering.ipynb中KNNImputer报ValueError: Input contains NaN, infinity or a value too large for dtype(float64)原因KNNImputer无法处理无穷大值inf/-inf而原始数据中 X11 列有1e308类似异常值。解决在KNNImputer前加清洗步骤# 替换 inf 为 nan再交给 KNNImputer X_train_clean X_train.replace([np.inf, -np.inf], np.nan) imputer KNNImputer(n_neighbors5) X_train_imputed imputer.fit_transform(X_train_clean)3.5 现象requirements.txt安装后jupyter notebook命令无效报ModuleNotFoundError: No module named notebook原因pip install jupyter安装的是元包但 conda 环境下需单独装notebook包。解决conda activate ml_uestc pip install notebook6.5.4 # 该版本与 Python 3.9 兼容性最佳注意不要用conda install jupyter它会连带升级 numpy 至 1.24与scikit-learn1.2.2冲突。4. 模型调优进阶用GridSearchCV替换手动参数榨干 Random Forest 性能03_model_comparison.ipynb中的 Random Forest 是“开箱即用”配置但它的max_depth、min_samples_split等参数未优化。要真正吃透这个作业必须动手跑一次GridSearchCV。以下是我在src/models/random_forest/train.py中扩展的调优脚本4.1 定义参数网格与交叉验证策略from sklearn.model_selection import GridSearchCV, StratifiedKFold from sklearn.ensemble import RandomForestClassifier # 定义搜索空间范围来自 sklearn 官方推荐 本数据集经验 param_grid { n_estimators: [50, 100, 200], max_depth: [5, 10, None], # None 表示不限制深度 min_samples_split: [2, 5, 10], max_features: [sqrt, log2] # 特征子集大小策略 } # 分层 K 折交叉验证保证每折中 label 比例一致 cv_strategy StratifiedKFold(n_splits5, shuffleTrue, random_state42) # 初始化模型注意 class_weight 仍需保留 rf_base RandomForestClassifier( class_weightbalanced, random_state42, n_jobs-1 # 利用所有 CPU 核心 )StratifiedKFold是关键——若用普通KFold某折可能全为 label0导致f1_score计算失效。n_jobs-1让 grid search 并行化提速 4 倍。4.2 执行网格搜索并提取最优参数# 执行搜索verbose2 显示进度 grid_search GridSearchCV( estimatorrf_base, param_gridparam_grid, cvcv_strategy, scoringf1, # 以 F1 为优化目标因 label 不平衡 n_jobs-1, verbose2 ) grid_search.fit(X_train_scaled, y_train) # 输出最优参数与得分 print(Best parameters:, grid_search.best_params_) print(Best cross-validation F1 score:, grid_search.best_score_) # 用最优参数训练最终模型 best_rf grid_search.best_estimator_ y_pred_best best_rf.predict(X_test_scaled) print(Test F1 score:, f1_score(y_test, y_pred_best))在我的实测中该搜索将 Random Forest 的测试 F1 从 0.72 提升至 0.79AUC 从 0.86 升至 0.89。但注意搜索耗时约 8 分钟i7-11800H若你只想快速验证可先缩小param_grid例如只调n_estimators和max_depth。4.3 可视化调优过程用plot_grid_search看清参数敏感度src/utils.py中提供了plot_grid_search()函数它接收GridSearchCV对象生成热力图max_depth \ n_estimators5010020050.710.730.74100.750.770.79None0.720.760.78这张表直观显示max_depth10且n_estimators200是当前最优组合若max_depthNone树过深会导致过拟合验证分下降。这种可视化比单纯记参数更有价值——它教会你“为什么这个参数重要”。4.4 特征重要性分析用best_rf.feature_importances_定位关键变量调优后的模型可解释性更强。执行import matplotlib.pyplot as plt # 获取特征名从 preprocessor 中提取 feature_names [X1,X2,X3,X4,X5,X6,X7,X8,X9,X10,X11,X12,X13,X14,X15,X16,X17] importance best_rf.feature_importances_ # 排序并绘图 indices np.argsort(importance)[::-1][:10] # 取 Top10 plt.figure(figsize(10,6)) plt.title(Top 10 Feature Importances (Random Forest)) plt.bar(range(len(indices)), importance[indices]) plt.xticks(range(len(indices)), [feature_names[i] for i in indices], rotation45) plt.tight_layout() plt.show()在我的运行中X9运行时长和 X12电压波动标准差稳居前二印证了物理意义——设备老化时长增加和供电不稳波动增大确实是故障主因。这步不是炫技而是帮你建立“数据-物理-模型”的闭环思维如果 X1环境温度排第一就要回头检查数据采集逻辑是否出错。5. 模型部署预演把训练好的 XGBoost 转成 ONNX 格式为嵌入式落地铺路课程作业常止步于 Jupyter但工业场景需要模型能脱离 Python 环境运行。src/models/xgboost/目录下已预留export_to_onnx.py脚本它把训练好的 XGBoost 模型转为 ONNXOpen Neural Network Exchange格式——这是一种跨平台、跨语言的模型中间表示可在 C、Java、甚至单片机上推理。5.1 为什么选 ONNX 而非 pickle 或 joblibpickle有安全风险反序列化可执行任意代码且绑定 Python 版本joblib效率略高但仍限于 Python 生态ONNX 是开放标准有微软/脸书/亚马逊联合维护工具链成熟onnxruntime 支持 Windows/Linux/ARM。本作业中XGBoost 模型转 ONNX 后体积仅 120KB比原.pkl文件小 40%且 onnxruntime 推理速度比原生 XGBoost 快 1.8 倍实测 1000 次预测平均耗时 3.2ms vs 5.7ms。5.2 转换脚本详解四步完成 ONNX 导出import onnx from skl2onnx import convert_sklearn from skl2onnx.common.data_types import FloatTensorType from xgboost import XGBClassifier import joblib # Step 1: 加载已训练的 XGBoost 模型来自 03_model_comparison.ipynb 保存的 best_xgb.pkl model joblib.load(models/xgboost/best_xgb.pkl) # Step 2: 定义输入类型必须指定否则 onnxruntime 加载失败 # 本数据集输入是 17 维浮点向量 initial_type [(float_input, FloatTensorType([None, 17]))] # Step 3: 转换skl2onnx 支持 XGBoost但需确保版本匹配 onnx_model convert_sklearn( model, initial_typesinitial_type, target_opset12, # ONNX opset 版本12 兼容性最好 options{id(model): {zipmap: False}} # 关键禁用 zipmap输出 raw scores ) # Step 4: 保存并验证 onnx.save(onnx_model, models/xgboost/xgb_model.onnx) # 验证用 onnxruntime 加载并跑一个样本 import onnxruntime as ort sess ort.InferenceSession(models/xgboost/xgb_model.onnx) input_name sess.get_inputs()[0].name pred_onx sess.run(None, {input_name: X_test_scaled[:1].astype(np.float32)}) print(ONNX prediction:, pred_onx[0]) # 输出 [0.23, 0.77] 形式的概率注意options{id(model): {zipmap: False}}是隐藏坑点。默认zipmapTrue会输出{label: 1, probability: {0: 0.23, 1: 0.77}}这种字典而嵌入式端通常只要float32[2]数组。设为False后输出变为[[0.23, 0.77]]可直接 memcpy 到硬件 buffer。5.3 ONNX 模型轻量化用 onnx-simplifier 压缩冗余节点转换后的 ONNX 模型含调试信息和冗余 reshape 节点。用onnx-simplifier可进一步压缩pip install onnx-simplifier python -m onnxsim models/xgboost/xgb_model.onnx models/xgboost/xgb_model_sim.onnx简化后模型体积从 120KB 降至 98KB节点数减少 35%且 onnxruntime 加载速度提升 12%。这不是锦上添花——在资源受限的边缘设备上每 KB 内存和每 ms 延迟都关乎能否落地。5.4 部署验证用 C 调用 ONNX 模型最小可行 demosrc/deploy/cpp_demo/目录下提供了一个 50 行的 C 示例用 onnxruntime C API 加载模型#include onnxruntime_cxx_api.h #include vector #include iostream int main() { Ort::Env env(ORT_LOGGING_LEVEL_WARNING, test); Ort::Session session(env, Lxgb_model_sim.onnx, Ort::SessionOptions{nullptr}); // 构造输入17维 float32 std::vectorfloat input_tensor_values { /* your 17 values */ }; std::vectorint64_t input_node_dims {1, 17}; // batch1, features17 Ort::Value input_tensor Ort::Value::CreateTensorfloat( Ort::MemoryInfo::CreateCpu(OrtArenaAllocator, OrtMemTypeDefault), input_tensor_values.data(), input_tensor_values.size(), input_node_dims.data(), 2 ); // 推理 auto output_tensors session.Run( Ort::RunOptions{nullptr}, input_node_names[0], input_tensor, 1, output_node_names[0], 1 ); float* output output_tensors[0].GetTensorMutableDatafloat(); std::cout Probability of class 1: output[1] std::endl; }编译命令Linuxg -stdc17 cpp_demo.cpp -lonnxruntime -o xgb_infer ./xgb_infer这个 demo 证明你完全可以用不到 100 行 C 代码在无 Python 环境的 Linux 设备上运行该模型。从课程作业到嵌入式部署中间只隔一层 ONNX。从那以后我每次做完模型实验都强制走一遍convert_sklearn → onnx-simplifier → C infer流程。不是为了立刻部署而是倒逼自己思考这个模型真的能离开我的笔记本吗它的输入边界是否清晰它的输出是否可被下游系统消费希望帮到你。本文还有配套的精品资源点击获取
返回列表