Make on-device OCR a pluggable local service so it runs locally on every platform (not just Windows), aimed at GoodNotes/Notability-class handwriting on low-power hardware (e.g. Zen2 APU, CPU/iGPU). - New OcrBackend abstraction (lib/services/ocr/): selector prefers an embedded ONNX recognition backend, falling back to the OS-native backend (Windows WinRT), and to a clean no-op when neither is available. - OnnxRecognitionBackend: flutter_onnxruntime session from a bundled asset, dart:ui preprocessing (resize to 48px, CHW float32, normalized), pure-Dart CTC greedy decode. Fully guarded — absent model/dict is a no-op; never throws. - ocr_engine.dart kept as a thin facade (recognizeImage) delegating to the selector, so ocr_service.dart is unchanged. - CtcDecoder unit-tested (6 tests). flutter analyze clean; all tests pass. - Model is not committed; tool/fetch_ocr_model.sh + assets/models/ocr/README.md document fetching PP-OCRv4 rec + dict on the dev machine. - CI: forward HTTPS_PROXY to the Windows build so CMake can fetch the ONNX Runtime native lib behind the GFW; README documents the system-install alternative. PP-OCR geometry/blank assumptions documented for on-device tuning. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2.2 KiB
Embedded OCR model (not committed)
The handwriting/text recognition backend
(lib/services/ocr/onnx_recognition_backend.dart) loads an ONNX recognition
model and its character dictionary from assets:
rec.onnx— the PP-OCRv4 mobile text recognition model (CTC, input3 x 48 x W, blank class index 0).ppocr_keys_v1.txt— the PP-OCR character dictionary, one character per line.
Neither file is committed to the repository (the model is large and the
dictionary is distributed with PaddleOCR). The app is built to treat their
absence as a clean no-op: if the model or dictionary is missing, the ONNX
backend reports unavailable and OCR falls back to the native platform backend
(or returns nothing). Only the .gitkeep placeholder is committed so the
assets/models/ocr/ asset directory is valid at build time.
How to obtain and place the files
Run the helper script on your development machine (it must download from the PaddleOCR sources and convert the Paddle inference model to ONNX):
./tool/fetch_ocr_model.sh
This places the two files here as:
assets/models/ocr/rec.onnx
assets/models/ocr/ppocr_keys_v1.txt
Sources
- PP-OCRv4 mobile recognition model (PaddleOCR inference model):
https://paddleocr.bj.bcebos.com/PP-OCRv4/chinese/ch_PP-OCRv4_rec_infer.tar
(English-only variant:
en_PP-OCRv4_rec_infer.tar) - Character dictionary
ppocr_keys_v1.txt: https://raw.githubusercontent.com/PaddlePaddle/PaddleOCR/main/ppocr/utils/ppocr_keys_v1.txt
Conversion
PaddleOCR ships Paddle inference models; convert to ONNX with paddle2onnx:
paddle2onnx \
--model_dir ch_PP-OCRv4_rec_infer \
--model_filename inference.pdmodel \
--params_filename inference.pdiparams \
--save_file rec.onnx \
--opset_version 14 \
--enable_onnx_checker True
Verification note
The backend assumes the PP-OCRv4 mobile rec convention (input 3 x 48 x W,
normalization (v/255 - 0.5)/0.5, CTC blank at index 0, dictionary shifted by
one). If you use a different exported model, verify the input shape,
normalization, and blank/dictionary convention and adjust
onnx_recognition_backend.dart / CtcDecoder accordingly.