VmF0x commited on
Commit
b47942e
·
verified ·
1 Parent(s): e9eefbb

Add Lapa Ukrainian HW-OCR LoRA adapter v1

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,119 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: lapa-llm/lapa-v0.1.2-instruct
3
+ library_name: peft
4
+ license: gemma
5
+ pipeline_tag: image-text-to-text
6
+ language:
7
+ - uk
8
+ tags:
9
+ - lora
10
+ - peft
11
+ - ocr
12
+ - handwriting
13
+ - handwritten-text-recognition
14
+ - ukrainian
15
+ - gemma3
16
+ - vision-language
17
+ ---
18
+
19
+ # Lapa Ukrainian Handwriting OCR — LoRA Adapter
20
+
21
+ LoRA adapter on top of [`lapa-llm/lapa-v0.1.2-instruct`](https://huggingface.co/lapa-llm/lapa-v0.1.2-instruct)
22
+ (a Gemma-3-12B Ukrainian vision-language model) for **Ukrainian handwritten-text
23
+ recognition (HTR / OCR)** on document crops.
24
+
25
+ The base Lapa model, applied zero-shot to handwriting crops, tends to *paraphrase*
26
+ rather than transcribe literally. This adapter retrains the text decoder to emit a
27
+ literal transcription of the text in the image. It was developed as an OCR component
28
+ for a Ukrainian HTR pipeline (handwritten + printed regions, math formulas).
29
+
30
+ ## Results (internal validation)
31
+
32
+ | Metric | Base Lapa (bf16) | + this LoRA |
33
+ |---|---|---|
34
+ | Handwritten CER | 3.28 | **0.113** |
35
+ | Handwritten exact-match | 1.3% | **47.7%** |
36
+ | Printed CER | 1.08 | **0.187** |
37
+
38
+ CER > 1 on the base reflects heavy paraphrasing (output far longer than ground truth).
39
+ The adapter removes that behavior and produces faithful transcriptions.
40
+
41
+ ## Intended use
42
+
43
+ - Transcribing **Ukrainian handwritten / printed text crops** (region-level images,
44
+ not full pages) into plain text.
45
+ - As a cross-vote / ensemble OCR partner alongside other VLMs.
46
+
47
+ Not tuned for: full-page layout, non-Ukrainian scripts, or marginal / very low-quality
48
+ regions (CER rises to ~0.55 on hard, low-confidence crops).
49
+
50
+ ## How to use
51
+
52
+ ```python
53
+ import torch
54
+ from PIL import Image
55
+ from peft import PeftModel
56
+ from transformers import AutoModelForImageTextToText, AutoProcessor
57
+
58
+ BASE = "lapa-llm/lapa-v0.1.2-instruct"
59
+ ADAPTER = "lapa-llm/lapa-ocr-lora" # this repo
60
+
61
+ base = AutoModelForImageTextToText.from_pretrained(
62
+ BASE,
63
+ torch_dtype=torch.bfloat16,
64
+ device_map="auto",
65
+ attn_implementation="sdpa",
66
+ )
67
+ model = PeftModel.from_pretrained(base, ADAPTER).eval()
68
+ processor = AutoProcessor.from_pretrained(BASE)
69
+
70
+ PROMPT = "Transcribe Ukrainian text literally. Output only the text, no preamble."
71
+ img = Image.open("crop.png").convert("RGB")
72
+ messages = [{
73
+ "role": "user",
74
+ "content": [
75
+ {"type": "image", "image": img},
76
+ {"type": "text", "text": PROMPT},
77
+ ],
78
+ }]
79
+
80
+ inputs = processor.apply_chat_template(
81
+ messages, add_generation_prompt=True, tokenize=True,
82
+ return_dict=True, return_tensors="pt", padding=True,
83
+ ).to(model.device, dtype=torch.bfloat16)
84
+
85
+ with torch.inference_mode():
86
+ gen = model.generate(**inputs, max_new_tokens=256, do_sample=False, num_beams=1)
87
+ text = processor.batch_decode(
88
+ gen[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True
89
+ )[0].strip()
90
+ print(text)
91
+ ```
92
+
93
+ ## Training
94
+
95
+ - **Base:** `lapa-llm/lapa-v0.1.2-instruct` (vision tower frozen; text decoder adapted)
96
+ - **Method:** LoRA (PEFT) — r=64, alpha=128, dropout=0.05, bias=none
97
+ - **Target modules:** `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj`
98
+ - **Task type:** `CAUSAL_LM`
99
+ - **Epochs:** 5 · **LR:** 1e-4 · **batch:** 2 × grad-accum 4 · **max_seq_len:** 1024
100
+ - **Precision:** bf16 · **Hardware:** 1× H100 80GB
101
+ - **Data:** Ukrainian handwritten / printed text crops with literal transcriptions.
102
+
103
+ ## License
104
+
105
+ This adapter is a derivative of Gemma-3 (via Lapa) and is released under the
106
+ **[Gemma Terms of Use](https://ai.google.dev/gemma/terms)**. Use is subject to the
107
+ [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy).
108
+ You must comply with the base model's license; see
109
+ [`lapa-llm/lapa-v0.1.2-instruct`](https://huggingface.co/lapa-llm/lapa-v0.1.2-instruct).
110
+
111
+ ## Acknowledgements
112
+
113
+ Built on the [Lapa LLM](https://huggingface.co/lapa-llm) by the Ukrainian Catholic
114
+ University, AGH University of Krakow, Igor Sikorsky Kyiv Polytechnic Institute, and
115
+ Lviv Polytechnic. Base model: Gemma-3-12B (Google DeepMind).
116
+
117
+ ### Framework versions
118
+ - PEFT 0.19.1
119
+ - Transformers (Gemma-3 support: ≥ 4.50)
adapter_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "lapa-llm/lapa-v0.1.2-instruct",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 128,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.05,
22
+ "lora_ga_config": null,
23
+ "megatron_config": null,
24
+ "megatron_core": "megatron.core",
25
+ "modules_to_save": null,
26
+ "peft_type": "LORA",
27
+ "peft_version": "0.19.1",
28
+ "qalora_group_size": 16,
29
+ "r": 64,
30
+ "rank_pattern": {},
31
+ "revision": null,
32
+ "target_modules": [
33
+ "k_proj",
34
+ "gate_proj",
35
+ "down_proj",
36
+ "up_proj",
37
+ "q_proj",
38
+ "v_proj",
39
+ "o_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_bdlora": null,
45
+ "use_dora": false,
46
+ "use_qalora": false,
47
+ "use_rslora": false
48
+ }
adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c8deff408306a5226ab731092678672ea61194af4bbdd1beb5bb2d06cf58ea85
3
+ size 1095430088
chat_template.jinja ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {{ bos_token }}
2
+ {%- if messages[0]['role'] == 'system' -%}
3
+ {%- if messages[0]['content'] is string -%}
4
+ {%- set first_user_prefix = messages[0]['content'] + '
5
+
6
+ ' -%}
7
+ {%- else -%}
8
+ {%- set first_user_prefix = messages[0]['content'][0]['text'] + '
9
+
10
+ ' -%}
11
+ {%- endif -%}
12
+ {%- set loop_messages = messages[1:] -%}
13
+ {%- else -%}
14
+ {%- set first_user_prefix = "" -%}
15
+ {%- set loop_messages = messages -%}
16
+ {%- endif -%}
17
+ {%- for message in loop_messages -%}
18
+ {%- if (message['role'] == 'user') != (loop.index0 % 2 == 0) -%}
19
+ {{ raise_exception("Conversation roles must alternate user/assistant/user/assistant/...") }}
20
+ {%- endif -%}
21
+ {%- if (message['role'] == 'assistant') -%}
22
+ {%- set role = "model" -%}
23
+ {%- else -%}
24
+ {%- set role = message['role'] -%}
25
+ {%- endif -%}
26
+ {{ '<start_of_turn>' + role + '
27
+ ' + (first_user_prefix if loop.first else "") }}
28
+ {%- if message['content'] is string -%}
29
+ {{ message['content'] | trim }}
30
+ {%- elif message['content'] is iterable -%}
31
+ {%- for item in message['content'] -%}
32
+ {%- if item['type'] == 'image' -%}
33
+ {{ '<start_of_image>' }}
34
+ {%- elif item['type'] == 'text' -%}
35
+ {{ item['text'] | trim }}
36
+ {%- endif -%}
37
+ {%- endfor -%}
38
+ {%- else -%}
39
+ {{ raise_exception("Invalid content type") }}
40
+ {%- endif -%}
41
+ {{ '<end_of_turn>
42
+ ' }}
43
+ {%- endfor -%}
44
+ {%- if add_generation_prompt -%}
45
+ {{'<start_of_turn>model
46
+ '}}
47
+ {%- endif -%}
processor_config.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "image_processor": {
3
+ "do_convert_rgb": null,
4
+ "do_normalize": true,
5
+ "do_rescale": true,
6
+ "do_resize": true,
7
+ "image_mean": [
8
+ 0.5,
9
+ 0.5,
10
+ 0.5
11
+ ],
12
+ "image_processor_type": "Gemma3ImageProcessor",
13
+ "image_seq_length": 256,
14
+ "image_std": [
15
+ 0.5,
16
+ 0.5,
17
+ 0.5
18
+ ],
19
+ "resample": 2,
20
+ "rescale_factor": 0.00392156862745098,
21
+ "size": {
22
+ "height": 896,
23
+ "width": 896
24
+ }
25
+ },
26
+ "image_seq_length": 256,
27
+ "processor_class": "Gemma3Processor"
28
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:96c223ece326885174a741fcb28446f49757a4f491b1187cdeb312c13b98b1ba
3
+ size 37379975
tokenizer_config.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "boi_token": "<start_of_image>",
4
+ "bos_token": "<bos>",
5
+ "clean_up_tokenization_spaces": false,
6
+ "eoi_token": "<end_of_image>",
7
+ "eos_token": "<eos>",
8
+ "image_token": "<image_soft_token>",
9
+ "is_local": false,
10
+ "local_files_only": false,
11
+ "mask_token": "<mask>",
12
+ "model_max_length": 1000000000000000019884624838656,
13
+ "model_specific_special_tokens": {
14
+ "boi_token": "<start_of_image>",
15
+ "eoi_token": "<end_of_image>",
16
+ "image_token": "<image_soft_token>"
17
+ },
18
+ "pad_token": "<pad>",
19
+ "processor_class": "Gemma3Processor",
20
+ "sp_model_kwargs": null,
21
+ "spaces_between_special_tokens": false,
22
+ "tokenizer_class": "GemmaTokenizer",
23
+ "unk_token": "<unk>",
24
+ "use_default_system_prompt": false
25
+ }