348 lines
8.6 KiB
Markdown
Executable File
348 lines
8.6 KiB
Markdown
Executable File
# Hệ Thống Quản Lý Model - Model Manager
|
|
|
|
## Tổng quan
|
|
|
|
Hệ thống **Model Manager** cho phép vận hành và quản lý tất cả các loại models trong dự án Land Classification, bao gồm:
|
|
- XGBoost
|
|
- Random Forest
|
|
- Decision Tree
|
|
- SVM
|
|
- CNN (PyTorch)
|
|
- Các model khác
|
|
|
|
## Cấu trúc
|
|
|
|
### 1. Model Storage
|
|
```
|
|
model_train/
|
|
├── model_odc.joblib # Model file
|
|
├── model_xgboost_20251221_172351.joblib
|
|
├── model_xgboost_20251221_172351_info.json # Metadata
|
|
├── model_cnn_20251221_163841.joblib
|
|
└── model_cnn_20251221_163841_info.json
|
|
```
|
|
|
|
### 2. Metadata Format
|
|
Mỗi model đi kèm với file JSON chứa metadata:
|
|
|
|
```json
|
|
{
|
|
"timestamp": "2025-12-21T17:23:57.306042",
|
|
"data_source": "Microsoft Planetary Computer STAC",
|
|
"collections": ["sentinel-2-l2a", "sentinel-1-rtc"],
|
|
"features": ["NDVI_mean", "VH_dB_mean", "VV_dB_mean"],
|
|
"model_type": "xgboost",
|
|
"n_features": 3,
|
|
"n_classes": 7,
|
|
"test_accuracy": 0.578125,
|
|
"train_accuracy": 1.0,
|
|
"classification_report": {...},
|
|
"confusion_matrix": [...],
|
|
"bbox": [105.6, 9.3, 106.2, 9.8],
|
|
"time_range": "2023-03-01/2023-05-31",
|
|
"resolution": 20
|
|
}
|
|
```
|
|
|
|
## Sử dụng
|
|
|
|
### 1. Trong Python Code
|
|
|
|
#### List tất cả models
|
|
```python
|
|
from model_manager import get_model_manager
|
|
|
|
model_manager = get_model_manager()
|
|
models = model_manager.list_models()
|
|
|
|
for model in models:
|
|
print(f"{model['filename']} - {model['model_type']} - Accuracy: {model['test_accuracy']}")
|
|
```
|
|
|
|
#### Load model
|
|
```python
|
|
model, encoder, metadata = model_manager.load_model("model_xgboost_20251221_172351.joblib")
|
|
|
|
print(f"Model type: {metadata['model_type']}")
|
|
print(f"Required features: {metadata['features']}")
|
|
```
|
|
|
|
#### Save model mới
|
|
```python
|
|
metadata = {
|
|
"timestamp": datetime.now().isoformat(),
|
|
"model_type": "random_forest",
|
|
"features": ["NDVI_mean", "VH_dB_mean", "VV_dB_mean"],
|
|
"n_features": 3,
|
|
"n_classes": 7,
|
|
"test_accuracy": 0.85,
|
|
"train_accuracy": 0.95
|
|
}
|
|
|
|
model_manager.save_model(
|
|
model=trained_model,
|
|
metadata=metadata,
|
|
model_filename="my_model.joblib",
|
|
label_encoder=encoder
|
|
)
|
|
```
|
|
|
|
#### Validate model
|
|
```python
|
|
validation = model_manager.validate_model("model_odc.joblib")
|
|
print(f"Valid: {validation['valid']}")
|
|
print(f"Errors: {validation['errors']}")
|
|
print(f"Warnings: {validation['warnings']}")
|
|
```
|
|
|
|
#### Get required features
|
|
```python
|
|
features = model_manager.get_required_features("model_xgboost_20251221_172351.joblib")
|
|
print(f"Required features: {features}")
|
|
```
|
|
|
|
### 2. Trong Notebook Training
|
|
|
|
File `01.train_ODC.ipynb` hoặc các notebook khác:
|
|
|
|
```python
|
|
# Import
|
|
from new_import_ODC import save_model
|
|
|
|
# Train model
|
|
model = RandomForestClassifier(n_estimators=100)
|
|
model.fit(X_train, y_train)
|
|
|
|
# Prepare metadata
|
|
metadata = {
|
|
"timestamp": datetime.now().isoformat(),
|
|
"model_type": "random_forest",
|
|
"features": ["ndvi"], # Danh sách features đã dùng
|
|
"n_features": 1,
|
|
"n_classes": len(np.unique(y_train)),
|
|
"test_accuracy": accuracy_score(y_test, y_pred),
|
|
"train_accuracy": model.score(X_train, y_train),
|
|
"data_source": "Local S3 ODC",
|
|
"training_samples": len(X_train),
|
|
"testing_samples": len(X_test)
|
|
}
|
|
|
|
# Save với metadata
|
|
save_model("model_odc.joblib", model, metadata=metadata, label_encoder=None)
|
|
```
|
|
|
|
### 3. Qua API
|
|
|
|
#### List models
|
|
```bash
|
|
curl http://localhost:8000/api/models/list
|
|
```
|
|
|
|
Response:
|
|
```json
|
|
{
|
|
"success": true,
|
|
"models": [
|
|
{
|
|
"filename": "model_xgboost_20251221_172351.joblib",
|
|
"model_type": "xgboost",
|
|
"features": ["NDVI_mean", "VH_dB_mean", "VV_dB_mean"],
|
|
"test_accuracy": 0.578125,
|
|
"size_mb": 0.45
|
|
}
|
|
],
|
|
"count": 3
|
|
}
|
|
```
|
|
|
|
#### Get model info
|
|
```bash
|
|
curl http://localhost:8000/api/models/model_odc.joblib/info
|
|
```
|
|
|
|
#### Validate model
|
|
```bash
|
|
curl http://localhost:8000/api/models/model_odc.joblib/validate
|
|
```
|
|
|
|
#### Delete model
|
|
```bash
|
|
curl -X DELETE http://localhost:8000/api/models/old_model.joblib
|
|
```
|
|
|
|
#### Predict với model cụ thể
|
|
```bash
|
|
curl -X POST http://localhost:8000/api/predict \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model_filename": "model_xgboost_20251221_172351.joblib",
|
|
"min_lon": 105.6,
|
|
"min_lat": 9.3,
|
|
"max_lon": 106.2,
|
|
"max_lat": 9.8,
|
|
"start_date": "2023-03-01",
|
|
"end_date": "2023-05-31"
|
|
}'
|
|
```
|
|
|
|
## Features Chính
|
|
|
|
### 1. Automatic Feature Detection
|
|
Hệ thống tự động detect features cần thiết từ metadata:
|
|
```python
|
|
metadata = model_manager._load_metadata("model.joblib")
|
|
required_features = metadata.get("features", [])
|
|
```
|
|
|
|
### 2. Model Type Support
|
|
Hỗ trợ nhiều loại model:
|
|
- **XGBoost**: GPU-accelerated gradient boosting
|
|
- **Random Forest**: Ensemble learning
|
|
- **Decision Tree**: Simple tree-based
|
|
- **SVM**: Support Vector Machine
|
|
- **CNN**: PyTorch neural networks
|
|
|
|
### 3. Backward Compatibility
|
|
Hệ thống vẫn hỗ trợ models cũ không có metadata:
|
|
- Tự động detect và tạo default metadata
|
|
- Load được cả format cũ (model only) và mới (dict với encoder)
|
|
|
|
### 4. Validation
|
|
Kiểm tra tính hợp lệ của model:
|
|
- File tồn tại
|
|
- Load được
|
|
- Metadata đầy đủ
|
|
- Features requirements
|
|
|
|
## Testing
|
|
|
|
Chạy test suite:
|
|
```bash
|
|
python test_model_manager.py
|
|
```
|
|
|
|
Output mẫu:
|
|
```
|
|
======================================================================
|
|
MODEL MANAGER TEST
|
|
======================================================================
|
|
|
|
✅ ModelManager initialized
|
|
|
|
======================================================================
|
|
TEST 1: LIST ALL MODELS
|
|
======================================================================
|
|
|
|
📦 Found 3 models:
|
|
|
|
[1] model_xgboost_20251221_172351.joblib
|
|
Size: 0.45 MB
|
|
Type: xgboost
|
|
Features: 3
|
|
Accuracy: 0.578125
|
|
|
|
[2] model_cnn_20251221_163841.joblib
|
|
Size: 0.12 MB
|
|
Type: cnn
|
|
Features: 3
|
|
Accuracy: 0.507812
|
|
|
|
[3] model_odc.joblib
|
|
Size: 0.02 MB
|
|
⚠️ No metadata
|
|
```
|
|
|
|
## Migration Guide
|
|
|
|
### Cho Models Cũ
|
|
|
|
Nếu bạn có models cũ không có metadata, có 2 cách:
|
|
|
|
#### Option 1: Tự động (Recommended)
|
|
Hệ thống sẽ tự động tạo default metadata khi load
|
|
|
|
#### Option 2: Tạo metadata manually
|
|
```python
|
|
# Tạo metadata file
|
|
metadata = {
|
|
"timestamp": "2025-12-21T12:00:00",
|
|
"model_type": "random_forest", # hoặc model type tương ứng
|
|
"features": ["ndvi"], # Features đã dùng khi train
|
|
"n_features": 1,
|
|
"n_classes": 8,
|
|
"test_accuracy": 0.75, # Nếu biết
|
|
}
|
|
|
|
import json
|
|
with open("model_train/model_odc_info.json", "w") as f:
|
|
json.dump(metadata, f, indent=2)
|
|
```
|
|
|
|
### Cho Training Code Mới
|
|
|
|
Luôn save model với metadata:
|
|
```python
|
|
save_model(
|
|
name_file="my_model.joblib",
|
|
model=trained_model,
|
|
metadata={...}, # Bắt buộc
|
|
label_encoder=encoder
|
|
)
|
|
```
|
|
|
|
## Best Practices
|
|
|
|
1. **Luôn include metadata** khi save model mới
|
|
2. **Sử dụng naming convention**: `model_{type}_{timestamp}.joblib`
|
|
3. **Test model** sau khi train: `model_manager.validate_model()`
|
|
4. **Document features** trong metadata để dễ sử dụng sau này
|
|
5. **Backup models** quan trọng trước khi xóa
|
|
|
|
## Troubleshooting
|
|
|
|
### Model không load được
|
|
```python
|
|
validation = model_manager.validate_model("model.joblib")
|
|
print(validation['errors']) # Xem lỗi cụ thể
|
|
```
|
|
|
|
### Thiếu metadata
|
|
Tạo metadata file manually (xem Migration Guide)
|
|
|
|
### Features không khớp
|
|
Kiểm tra `metadata['features']` và đảm bảo data đầu vào có đúng features
|
|
|
|
## API Endpoints Summary
|
|
|
|
| Endpoint | Method | Description |
|
|
|----------|--------|-------------|
|
|
| `/api/models/list` | GET | List all models |
|
|
| `/api/models/{filename}/info` | GET | Get model details |
|
|
| `/api/models/{filename}/validate` | GET | Validate model |
|
|
| `/api/models/{filename}` | DELETE | Delete model |
|
|
| `/api/predict` | POST | Predict with model |
|
|
| `/api/batch/predict` | POST | Batch prediction |
|
|
| `/api/predict-with-ndvi` | POST | Predict + NDVI export |
|
|
|
|
## File Structure
|
|
|
|
```
|
|
remote-sensing/
|
|
├── model_manager.py # Core ModelManager class
|
|
├── test_model_manager.py # Test suite
|
|
├── new_import_ODC.py # Updated save_model function
|
|
├── train_module.py # Updated training module
|
|
├── api_server.py # API với ModelManager integration
|
|
└── model_train/ # Models directory
|
|
├── *.joblib # Model files
|
|
└── *_info.json # Metadata files
|
|
```
|
|
|
|
## Next Steps
|
|
|
|
1. ✅ Migrate existing notebooks để sử dụng metadata
|
|
2. ✅ Update UI để cho phép chọn model
|
|
3. ✅ Add model comparison features
|
|
4. ✅ Implement model versioning
|
|
5. ✅ Add automated model backup
|