mirror of
https://git.victorphan.net/basketballcantho/book_translator.git
synced 2026-08-05 06:33:11 +07:00
first commit
This commit is contained in:
@@ -0,0 +1,101 @@
|
||||
# 📚 Bilingual EPUB Translator (English → English/Vietnamese)
|
||||
|
||||
Translate any English `.epub` book into a beautifully formatted bilingual English-Vietnamese edition. The LLM runs **directly in Python** via HuggingFace Transformers — no Ollama or external server needed.
|
||||
|
||||
The model downloads once from HuggingFace Hub and is cached at `~/.cache/huggingface` for all future runs.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- **Python 3.10+** (conda env recommended)
|
||||
- **256 GB RAM** (for qwen2.5:72b in INT8 quantization)
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
conda activate env03
|
||||
cd /home/x79/book_translator
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
> **Note:** The first run will download the model weights (~72 GB for INT8). This is a one-time download cached at `~/.cache/huggingface`.
|
||||
|
||||
## Usage
|
||||
|
||||
### Basic (uses defaults)
|
||||
|
||||
```bash
|
||||
python translate_epub.py "Mona___s_Eyes___40_Thomas_Schlesser__Hildegarde_Serle___40_translator__41___41_.epub"
|
||||
```
|
||||
|
||||
This produces `Mona___s_Eyes_..._bilingual.epub` in the same directory.
|
||||
|
||||
### Specify output path
|
||||
|
||||
```bash
|
||||
python translate_epub.py input.epub -o my_bilingual_book.epub
|
||||
```
|
||||
|
||||
### Use a different / smaller model
|
||||
|
||||
```bash
|
||||
# Smaller model for faster inference
|
||||
python translate_epub.py input.epub --model Qwen/Qwen2.5-7B-Instruct
|
||||
|
||||
# Another model entirely
|
||||
python translate_epub.py input.epub --model meta-llama/Llama-3-8B-Instruct
|
||||
```
|
||||
|
||||
### Disable INT8 quantization (uses more RAM)
|
||||
|
||||
```bash
|
||||
python translate_epub.py input.epub --no-8bit
|
||||
```
|
||||
|
||||
## All Options
|
||||
|
||||
| Flag | Default | Description |
|
||||
|------|---------|-------------|
|
||||
| `input` | *(required)* | Path to the English `.epub` file |
|
||||
| `-o, --output` | `<input>_bilingual.epub` | Output file path |
|
||||
| `--model` | `Qwen/Qwen2.5-72B-Instruct` | HuggingFace model ID |
|
||||
| `--chunk-size` | `5` | Paragraphs per inference call |
|
||||
| `--max-new-tokens` | `8192` | Max generated tokens per chunk |
|
||||
| `--no-8bit` | off | Disable INT8 quantization |
|
||||
| `--no-resume` | off | Ignore cached progress, re-translate all |
|
||||
| `--clear-cache` | — | Delete cached checkpoints and exit |
|
||||
|
||||
## Checkpoint / Resume
|
||||
|
||||
Since CPU inference on a 72B model can take **many hours**, the script saves each chapter's translated HTML to a `.bilingual_cache_*` directory next to your input file.
|
||||
|
||||
If the script crashes or you `Ctrl+C`:
|
||||
|
||||
```bash
|
||||
# Just re-run the same command — it picks up where it left off
|
||||
python translate_epub.py input.epub
|
||||
```
|
||||
|
||||
To force a fresh start:
|
||||
|
||||
```bash
|
||||
python translate_epub.py input.epub --no-resume
|
||||
# or delete the cache entirely:
|
||||
python translate_epub.py input.epub --clear-cache
|
||||
```
|
||||
|
||||
## Output Format
|
||||
|
||||
The bilingual EPUB contains, for every paragraph:
|
||||
|
||||
1. **English** (bold) — the original text
|
||||
2. *Vietnamese* (italic, red left border) — the translation
|
||||
3. A **Vocabulary Box** at the end of each chunk with 3–5 difficult words explained
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Problem | Solution |
|
||||
|---------|----------|
|
||||
| Out of memory | Use a smaller model: `--model Qwen/Qwen2.5-7B-Instruct` |
|
||||
| Slow inference | Reduce chunk size: `--chunk-size 3` |
|
||||
| Model download fails | Check disk space at `~/.cache/huggingface` and internet |
|
||||
| Poor translation | Try `--no-8bit` for full precision (needs more RAM) |
|
||||
Reference in New Issue
Block a user