# Hệ Thống Quản Lý Model - Model Manager ## Tổng quan Hệ thống **Model Manager** cho phép vận hành và quản lý tất cả các loại models trong dự án Land Classification, bao gồm: - XGBoost - Random Forest - Decision Tree - SVM - CNN (PyTorch) - Các model khác ## Cấu trúc ### 1. Model Storage ``` model_train/ ├── model_odc.joblib # Model file ├── model_xgboost_20251221_172351.joblib ├── model_xgboost_20251221_172351_info.json # Metadata ├── model_cnn_20251221_163841.joblib └── model_cnn_20251221_163841_info.json ``` ### 2. Metadata Format Mỗi model đi kèm với file JSON chứa metadata: ```json { "timestamp": "2025-12-21T17:23:57.306042", "data_source": "Microsoft Planetary Computer STAC", "collections": ["sentinel-2-l2a", "sentinel-1-rtc"], "features": ["NDVI_mean", "VH_dB_mean", "VV_dB_mean"], "model_type": "xgboost", "n_features": 3, "n_classes": 7, "test_accuracy": 0.578125, "train_accuracy": 1.0, "classification_report": {...}, "confusion_matrix": [...], "bbox": [105.6, 9.3, 106.2, 9.8], "time_range": "2023-03-01/2023-05-31", "resolution": 20 } ``` ## Sử dụng ### 1. Trong Python Code #### List tất cả models ```python from model_manager import get_model_manager model_manager = get_model_manager() models = model_manager.list_models() for model in models: print(f"{model['filename']} - {model['model_type']} - Accuracy: {model['test_accuracy']}") ``` #### Load model ```python model, encoder, metadata = model_manager.load_model("model_xgboost_20251221_172351.joblib") print(f"Model type: {metadata['model_type']}") print(f"Required features: {metadata['features']}") ``` #### Save model mới ```python metadata = { "timestamp": datetime.now().isoformat(), "model_type": "random_forest", "features": ["NDVI_mean", "VH_dB_mean", "VV_dB_mean"], "n_features": 3, "n_classes": 7, "test_accuracy": 0.85, "train_accuracy": 0.95 } model_manager.save_model( model=trained_model, metadata=metadata, model_filename="my_model.joblib", label_encoder=encoder ) ``` #### Validate model ```python validation = model_manager.validate_model("model_odc.joblib") print(f"Valid: {validation['valid']}") print(f"Errors: {validation['errors']}") print(f"Warnings: {validation['warnings']}") ``` #### Get required features ```python features = model_manager.get_required_features("model_xgboost_20251221_172351.joblib") print(f"Required features: {features}") ``` ### 2. Trong Notebook Training File `01.train_ODC.ipynb` hoặc các notebook khác: ```python # Import from new_import_ODC import save_model # Train model model = RandomForestClassifier(n_estimators=100) model.fit(X_train, y_train) # Prepare metadata metadata = { "timestamp": datetime.now().isoformat(), "model_type": "random_forest", "features": ["ndvi"], # Danh sách features đã dùng "n_features": 1, "n_classes": len(np.unique(y_train)), "test_accuracy": accuracy_score(y_test, y_pred), "train_accuracy": model.score(X_train, y_train), "data_source": "Local S3 ODC", "training_samples": len(X_train), "testing_samples": len(X_test) } # Save với metadata save_model("model_odc.joblib", model, metadata=metadata, label_encoder=None) ``` ### 3. Qua API #### List models ```bash curl http://localhost:8000/api/models/list ``` Response: ```json { "success": true, "models": [ { "filename": "model_xgboost_20251221_172351.joblib", "model_type": "xgboost", "features": ["NDVI_mean", "VH_dB_mean", "VV_dB_mean"], "test_accuracy": 0.578125, "size_mb": 0.45 } ], "count": 3 } ``` #### Get model info ```bash curl http://localhost:8000/api/models/model_odc.joblib/info ``` #### Validate model ```bash curl http://localhost:8000/api/models/model_odc.joblib/validate ``` #### Delete model ```bash curl -X DELETE http://localhost:8000/api/models/old_model.joblib ``` #### Predict với model cụ thể ```bash curl -X POST http://localhost:8000/api/predict \ -H "Content-Type: application/json" \ -d '{ "model_filename": "model_xgboost_20251221_172351.joblib", "min_lon": 105.6, "min_lat": 9.3, "max_lon": 106.2, "max_lat": 9.8, "start_date": "2023-03-01", "end_date": "2023-05-31" }' ``` ## Features Chính ### 1. Automatic Feature Detection Hệ thống tự động detect features cần thiết từ metadata: ```python metadata = model_manager._load_metadata("model.joblib") required_features = metadata.get("features", []) ``` ### 2. Model Type Support Hỗ trợ nhiều loại model: - **XGBoost**: GPU-accelerated gradient boosting - **Random Forest**: Ensemble learning - **Decision Tree**: Simple tree-based - **SVM**: Support Vector Machine - **CNN**: PyTorch neural networks ### 3. Backward Compatibility Hệ thống vẫn hỗ trợ models cũ không có metadata: - Tự động detect và tạo default metadata - Load được cả format cũ (model only) và mới (dict với encoder) ### 4. Validation Kiểm tra tính hợp lệ của model: - File tồn tại - Load được - Metadata đầy đủ - Features requirements ## Testing Chạy test suite: ```bash python test_model_manager.py ``` Output mẫu: ``` ====================================================================== MODEL MANAGER TEST ====================================================================== ✅ ModelManager initialized ====================================================================== TEST 1: LIST ALL MODELS ====================================================================== 📦 Found 3 models: [1] model_xgboost_20251221_172351.joblib Size: 0.45 MB Type: xgboost Features: 3 Accuracy: 0.578125 [2] model_cnn_20251221_163841.joblib Size: 0.12 MB Type: cnn Features: 3 Accuracy: 0.507812 [3] model_odc.joblib Size: 0.02 MB ⚠️ No metadata ``` ## Migration Guide ### Cho Models Cũ Nếu bạn có models cũ không có metadata, có 2 cách: #### Option 1: Tự động (Recommended) Hệ thống sẽ tự động tạo default metadata khi load #### Option 2: Tạo metadata manually ```python # Tạo metadata file metadata = { "timestamp": "2025-12-21T12:00:00", "model_type": "random_forest", # hoặc model type tương ứng "features": ["ndvi"], # Features đã dùng khi train "n_features": 1, "n_classes": 8, "test_accuracy": 0.75, # Nếu biết } import json with open("model_train/model_odc_info.json", "w") as f: json.dump(metadata, f, indent=2) ``` ### Cho Training Code Mới Luôn save model với metadata: ```python save_model( name_file="my_model.joblib", model=trained_model, metadata={...}, # Bắt buộc label_encoder=encoder ) ``` ## Best Practices 1. **Luôn include metadata** khi save model mới 2. **Sử dụng naming convention**: `model_{type}_{timestamp}.joblib` 3. **Test model** sau khi train: `model_manager.validate_model()` 4. **Document features** trong metadata để dễ sử dụng sau này 5. **Backup models** quan trọng trước khi xóa ## Troubleshooting ### Model không load được ```python validation = model_manager.validate_model("model.joblib") print(validation['errors']) # Xem lỗi cụ thể ``` ### Thiếu metadata Tạo metadata file manually (xem Migration Guide) ### Features không khớp Kiểm tra `metadata['features']` và đảm bảo data đầu vào có đúng features ## API Endpoints Summary | Endpoint | Method | Description | |----------|--------|-------------| | `/api/models/list` | GET | List all models | | `/api/models/{filename}/info` | GET | Get model details | | `/api/models/{filename}/validate` | GET | Validate model | | `/api/models/{filename}` | DELETE | Delete model | | `/api/predict` | POST | Predict with model | | `/api/batch/predict` | POST | Batch prediction | | `/api/predict-with-ndvi` | POST | Predict + NDVI export | ## File Structure ``` remote-sensing/ ├── model_manager.py # Core ModelManager class ├── test_model_manager.py # Test suite ├── new_import_ODC.py # Updated save_model function ├── train_module.py # Updated training module ├── api_server.py # API với ModelManager integration └── model_train/ # Models directory ├── *.joblib # Model files └── *_info.json # Metadata files ``` ## Next Steps 1. ✅ Migrate existing notebooks để sử dụng metadata 2. ✅ Update UI để cho phép chọn model 3. ✅ Add model comparison features 4. ✅ Implement model versioning 5. ✅ Add automated model backup