Files
2026-06-04 14:10:37 +07:00

102 lines
3.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 📚 Bilingual EPUB Translator (English → English/Vietnamese)
Translate any English `.epub` book into a beautifully formatted bilingual English-Vietnamese edition. The LLM runs **directly in Python** via HuggingFace Transformers — no Ollama or external server needed.
The model downloads once from HuggingFace Hub and is cached at `~/.cache/huggingface` for all future runs.
## Prerequisites
- **Python 3.10+** (conda env recommended)
- **256 GB RAM** (for qwen2.5:72b in INT8 quantization)
## Installation
```bash
conda activate env03
cd /home/x79/book_translator
pip install -r requirements.txt
```
> **Note:** The first run will download the model weights (~72 GB for INT8). This is a one-time download cached at `~/.cache/huggingface`.
## Usage
### Basic (uses defaults)
```bash
python translate_epub.py "Mona___s_Eyes___40_Thomas_Schlesser__Hildegarde_Serle___40_translator__41___41_.epub"
```
This produces `Mona___s_Eyes_..._bilingual.epub` in the same directory.
### Specify output path
```bash
python translate_epub.py input.epub -o my_bilingual_book.epub
```
### Use a different / smaller model
```bash
# Smaller model for faster inference
python translate_epub.py input.epub --model Qwen/Qwen2.5-7B-Instruct
# Another model entirely
python translate_epub.py input.epub --model meta-llama/Llama-3-8B-Instruct
```
### Disable INT8 quantization (uses more RAM)
```bash
python translate_epub.py input.epub --no-8bit
```
## All Options
| Flag | Default | Description |
|------|---------|-------------|
| `input` | *(required)* | Path to the English `.epub` file |
| `-o, --output` | `<input>_bilingual.epub` | Output file path |
| `--model` | `Qwen/Qwen2.5-72B-Instruct` | HuggingFace model ID |
| `--chunk-size` | `5` | Paragraphs per inference call |
| `--max-new-tokens` | `8192` | Max generated tokens per chunk |
| `--no-8bit` | off | Disable INT8 quantization |
| `--no-resume` | off | Ignore cached progress, re-translate all |
| `--clear-cache` | — | Delete cached checkpoints and exit |
## Checkpoint / Resume
Since CPU inference on a 72B model can take **many hours**, the script saves each chapter's translated HTML to a `.bilingual_cache_*` directory next to your input file.
If the script crashes or you `Ctrl+C`:
```bash
# Just re-run the same command — it picks up where it left off
python translate_epub.py input.epub
```
To force a fresh start:
```bash
python translate_epub.py input.epub --no-resume
# or delete the cache entirely:
python translate_epub.py input.epub --clear-cache
```
## Output Format
The bilingual EPUB contains, for every paragraph:
1. **English** (bold) — the original text
2. *Vietnamese* (italic, red left border) — the translation
3. A **Vocabulary Box** at the end of each chunk with 35 difficult words explained
## Troubleshooting
| Problem | Solution |
|---------|----------|
| Out of memory | Use a smaller model: `--model Qwen/Qwen2.5-7B-Instruct` |
| Slow inference | Reduce chunk size: `--chunk-size 3` |
| Model download fails | Check disk space at `~/.cache/huggingface` and internet |
| Poor translation | Try `--no-8bit` for full precision (needs more RAM) |