OCR: embedded, cross-platform ONNX backend with pluggable fallback
Some checks failed
CI / Flutter (analyze, test, Windows build) (push) Failing after 30s
CI / Server tests (optional) (push) Failing after 29s

Make on-device OCR a pluggable local service so it runs locally on every
platform (not just Windows), aimed at GoodNotes/Notability-class handwriting on
low-power hardware (e.g. Zen2 APU, CPU/iGPU).

- New OcrBackend abstraction (lib/services/ocr/): selector prefers an embedded
  ONNX recognition backend, falling back to the OS-native backend (Windows
  WinRT), and to a clean no-op when neither is available.
- OnnxRecognitionBackend: flutter_onnxruntime session from a bundled asset,
  dart:ui preprocessing (resize to 48px, CHW float32, normalized), pure-Dart CTC
  greedy decode. Fully guarded — absent model/dict is a no-op; never throws.
- ocr_engine.dart kept as a thin facade (recognizeImage) delegating to the
  selector, so ocr_service.dart is unchanged.
- CtcDecoder unit-tested (6 tests). flutter analyze clean; all tests pass.
- Model is not committed; tool/fetch_ocr_model.sh + assets/models/ocr/README.md
  document fetching PP-OCRv4 rec + dict on the dev machine.
- CI: forward HTTPS_PROXY to the Windows build so CMake can fetch the ONNX
  Runtime native lib behind the GFW; README documents the system-install
  alternative. PP-OCR geometry/blank assumptions documented for on-device tuning.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-21 03:51:54 +08:00
parent 25ba717c97
commit 99b98b96b0
19 changed files with 684 additions and 18 deletions

View File

@@ -0,0 +1,63 @@
# Embedded OCR model (not committed)
The handwriting/text recognition backend
(`lib/services/ocr/onnx_recognition_backend.dart`) loads an ONNX recognition
model and its character dictionary **from assets**:
- `rec.onnx` — the PP-OCRv4 mobile text recognition model (CTC, input
`3 x 48 x W`, blank class index 0).
- `ppocr_keys_v1.txt` — the PP-OCR character dictionary, one character per line.
Neither file is committed to the repository (the model is large and the
dictionary is distributed with PaddleOCR). The app is built to treat their
absence as a clean no-op: if the model or dictionary is missing, the ONNX
backend reports unavailable and OCR falls back to the native platform backend
(or returns nothing). Only the `.gitkeep` placeholder is committed so the
`assets/models/ocr/` asset directory is valid at build time.
## How to obtain and place the files
Run the helper script on your development machine (it must download from the
PaddleOCR sources and convert the Paddle inference model to ONNX):
```bash
./tool/fetch_ocr_model.sh
```
This places the two files here as:
```
assets/models/ocr/rec.onnx
assets/models/ocr/ppocr_keys_v1.txt
```
### Sources
- PP-OCRv4 mobile recognition model (PaddleOCR inference model):
https://paddleocr.bj.bcebos.com/PP-OCRv4/chinese/ch_PP-OCRv4_rec_infer.tar
(English-only variant: `en_PP-OCRv4_rec_infer.tar`)
- Character dictionary `ppocr_keys_v1.txt`:
https://raw.githubusercontent.com/PaddlePaddle/PaddleOCR/main/ppocr/utils/ppocr_keys_v1.txt
### Conversion
PaddleOCR ships Paddle inference models; convert to ONNX with
[paddle2onnx](https://github.com/PaddlePaddle/Paddle2ONNX):
```bash
paddle2onnx \
--model_dir ch_PP-OCRv4_rec_infer \
--model_filename inference.pdmodel \
--params_filename inference.pdiparams \
--save_file rec.onnx \
--opset_version 14 \
--enable_onnx_checker True
```
## Verification note
The backend assumes the PP-OCRv4 mobile rec convention (input `3 x 48 x W`,
normalization `(v/255 - 0.5)/0.5`, CTC blank at index 0, dictionary shifted by
one). If you use a different exported model, verify the input shape,
normalization, and blank/dictionary convention and adjust
`onnx_recognition_backend.dart` / `CtcDecoder` accordingly.