Files
CSIROBoeingPhase5-Vietnam/QUICK_REFERENCE.md
T
Victor Phan 7db1027ac6 update 01
2025-11-12 13:21:43 +07:00

5.8 KiB

Quick Reference: 3-Step Workflow

3 Bước Đơn Giản

Step 1️⃣: SERVER - Load Data Raw

Notebook: 01.prepare_data_on_server.ipynb
Location: Run on server
Time: 10-20 min
Output: 2 files NetCDF (~80-100 GB)

What it does:

  • Tải S2 (red, nir, scl) từ S3
  • Tải S1 (VH, VV) từ S3
  • Lưu 2 file NetCDF thô (chưa xử lý)
  • Chép file training shapefile

How to run:

# Run cells in order (1-8)
# Watch for progress bars in cell 5: [01/13], [02/13], ..., [13/13]
# Expected: ✅ Success! Shape: {'time': 396, 'y': 10000, 'x': 10000}

Output:

data_for_training/
├─ sentinel2_raw.nc      (S2 thô, ~50 GB)
├─ sentinel1_raw.nc      (S1 thô, ~30 GB)
└─ train_data/
   ├─ *.shp, *.shx, *.dbf (training points)

Step 2️⃣: LOCAL - Process & Train

Notebook: 02.process_and_train_local.ipynb
Location: Download data + run on local machine
Time: 30-60 min (CPU) or 10-15 min (GPU)
Output: Trained PyTorch CNN model (~100 MB)

What it does:

  • Load NetCDF files
  • Cloud mask (SCL band)
  • Calculate NDVI
  • Fill missing values
  • Monthly aggregation
  • Extract features at training points
  • Train PyTorch CNN (50 epochs)
  • Evaluate on test set
  • Save model

How to run:

# Make sure data_for_training/ folder exists locally
# Run cells in order (1-11)
# Watch training progress: epoch 1/50, epoch 2/50, ...
# Expected: ✅ Test Accuracy: 0.7-0.85

Output:

model_cnn_pytorch_local.pth         (Model weights)
model_cnn_pytorch_local_checkpoint.pth (Full checkpoint)

Step 3️⃣: LOCAL - Make Predictions

Notebook: 03.predict_CNN_PyTorch_local.ipynb
Location: Run on local machine
Time: 5-10 min
Output: Classification maps (SHP/TIF)

What it does:

  • Load trained model
  • Process full spatial data
  • Apply model to every pixel
  • Generate classification map
  • Save as SHP/TIF format

How to run:

# Make sure model file exists locally
# Run cells in order
# Expected: ✅ Classification map generated with 8 classes

Output:

classification_map.shp          (Land use map)
classification_map.tif          (GeoTIFF format)

File Structure

On Server:
────────────
/server/path/01.prepare_data_on_server.ipynb
→ Outputs to: data_for_training/ (80-100 GB)


On Local Machine:
─────────────────
/local/path/
├─ data_for_training/         ← Downloaded from server
│  ├─ sentinel2_raw.nc
│  ├─ sentinel1_raw.nc
│  └─ train_data/
│
├─ 02.process_and_train_local.ipynb
├─ 03.predict_CNN_PyTorch_local.ipynb
│
├─ model_cnn_pytorch_local.pth         ← Generated by step 2
├─ model_cnn_pytorch_local_checkpoint.pth
│
└─ classification_map.shp              ← Generated by step 3

Key Differences from Old Workflow

Aspect Old New
Processing Server does everything Server loads, local processes
Speed Slow (server overloaded) Fast (parallel processing)
Memory 403 TB attempt (crash!) 20 GB (manageable)
Flexibility Hard to debug Easy to iterate locally
Re-processing Must go back to server Can redo locally anytime

Checklist

Before Step 1:

  • Server has Dask + Datacube + S3 access
  • At least 500 GB free on server
  • Network stable

Before Step 2:

  • Downloaded all data from server
  • At least 100 GB free on local machine
  • Local machine has Python + PyTorch installed
  • GPU available (optional but faster)

Before Step 3:

  • Notebook 02 completed with accuracy ≥ 0.70
  • Model file exists locally
  • Processed data available

Troubleshooting

Problem Solution
Server: "403 TB OOM" Already fixed! Using monthly chunking in cell 5
Server: "S3 access denied" Check credentials in cell 2
Local: "File not found" Verify data_for_training/ folder location
Local: "Model accuracy too low" Check cloud masking - increase training epochs
Local: "Out of memory" Close other apps, reduce batch_size in training

Performance Expectations

Step Task Time (CPU) Time (GPU)
1️⃣ Load (Server) S2 + S1 download 10-20 min N/A
🔄 Transfer Download to local 30-60 min 30-60 min
2️⃣ Process & Train (Local) All preprocessing + CNN 30-60 min 10-15 min
3️⃣ Predict (Local) Full spatial predictions 5-10 min 2-5 min
TOTAL All steps 1-2 hours 1-1.5 hours

Expected Results

After Step 1:

✅ 2 NetCDF files (80-100 GB)
✅ Training shapefile (1130 points)
✅ Ready to download

After Step 2:

✅ Model trained (50 epochs completed)
✅ Test accuracy: 70-85%
✅ All 8 classes learned
✅ Model saved (100 MB)

After Step 3:

✅ Classification map generated
✅ 8 classes distributed
✅ Accuracy reasonable on test areas
✅ Output in SHP/TIF format

Why This Design?

Server chỉ load (không xử lý):

  • Tránh lãng phí tài nguyên server
  • Tải nhanh, server sẵn cho task khác
  • Monthly chunking giải quyết OOM

Local chỉ xử lý (không load):

  • Toàn quyền kiểm soát quy trình
  • Dễ debug & iterate
  • GPU nếu có → nhanh

Kết quả:

  • Không bao giờ OOM
  • Tất cả hoạt động nhanh
  • Dễ tái tạo & tùy chỉnh

Next Steps

  1. Today: Run Notebook 01 on server
  2. Tonight: Download data (~30-60 min)
  3. Tomorrow: Run Notebook 02 (train model)
  4. Tomorrow: Run Notebook 03 (make predictions)
  5. Day after: Analyze results

Total timeline: 2-3 days (with overnight download)


Design: Server loads → Local processes
Status: Ready to use
Created: November 12, 2025