multimodalart HF Staff commited on
Commit
f99c6cc
·
verified ·
1 Parent(s): 8288c04

Add agent recipe (agent.md) for curl-based LoRA training

Browse files
Files changed (1) hide show
  1. agent.md +91 -0
agent.md ADDED
@@ -0,0 +1,91 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Train a Krea 2 LoRA — agent recipe
2
+
3
+ You can train a custom **Krea 2** LoRA on Hugging Face Jobs. Krea 2 weights are public, so you
4
+ only need the USER's own write-scoped HF token (https://huggingface.co/settings/tokens). Every
5
+ step below runs under that token and is billed to that account.
6
+
7
+ ## Inputs to collect from the user
8
+ - **lora_name** — a unique name. The result is pushed to `<user>/<lora_name>`.
9
+ - **trigger** — a style phrase (e.g. `"in the style of TOK"`) or, for a subject/face/object, a
10
+ rare token (e.g. `"TOK person"`). It is included in every training caption.
11
+ - **concept** — one of: `style` | `character` | `object` | `face`.
12
+ - **images** — 4–30 images. The user will give you ONE of:
13
+ (a) an existing HF dataset repo with `image` + `prompt` columns → use it directly, skip Step 1;
14
+ (b) a local folder of images → build the dataset (Step 1);
15
+ (c) a request like *"find impasto-style paintings"* → gather the images yourself first, then
16
+ build the dataset (Step 1).
17
+
18
+ ## Step 1 — build the dataset (skip if the user already has one)
19
+ Create an imagefolder dataset under the user's account; each row is an image + a caption that
20
+ contains the trigger.
21
+
22
+ ```python
23
+ import json, os
24
+ from pathlib import Path
25
+ from huggingface_hub import HfApi
26
+
27
+ TOKEN = os.environ["HF_TOKEN"] # the user's write token
28
+ REPO = "<user>/<lora_name>-dataset"
29
+ TRIGGER = "<trigger>"
30
+ images = ["img1.png", "img2.png"] # local paths you gathered
31
+ captions = ["a cat on a sofa, " + TRIGGER, "a cat by a window, " + TRIGGER] # 1 per image
32
+
33
+ api = HfApi(token=TOKEN)
34
+ api.create_repo(REPO, repo_type="dataset", private=True, exist_ok=True)
35
+ tmp = Path("ds"); tmp.mkdir(exist_ok=True)
36
+ rows = []
37
+ for i, (src, cap) in enumerate(zip(images, captions)):
38
+ fn = f"{i:03d}{Path(src).suffix.lower()}"
39
+ (tmp / fn).write_bytes(Path(src).read_bytes())
40
+ rows.append({"file_name": fn, "prompt": cap})
41
+ (tmp / "metadata.jsonl").write_text("\n".join(json.dumps(r) for r in rows))
42
+ api.upload_folder(repo_id=REPO, repo_type="dataset", folder_path=str(tmp))
43
+ print("dataset:", REPO)
44
+ ```
45
+
46
+ No captions yet? Use the trigger alone as every caption — fine for a single-concept style.
47
+
48
+ ## Step 2 — launch the training job
49
+ Download the ready-made job script and submit it on HF Jobs under the user's token:
50
+
51
+ ```bash
52
+ export HF_TOKEN=hf_... # the user's write token
53
+ curl -sO https://huggingface.co/spaces/multimodalart/krea2-lora-trainer/raw/main/train_job.py
54
+ hf jobs uv run --flavor rtx-pro-6000 --timeout 1h -s HF_TOKEN=$HF_TOKEN train_job.py -- \
55
+ --dataset <user>/<lora_name>-dataset \
56
+ --lora-name <lora_name> \
57
+ --trigger "<trigger>" \
58
+ --concept style
59
+ ```
60
+
61
+ Python equivalent:
62
+
63
+ ```python
64
+ from huggingface_hub import HfApi
65
+ api = HfApi(token=TOKEN)
66
+ job = api.run_uv_job(
67
+ "https://huggingface.co/spaces/multimodalart/krea2-lora-trainer/raw/main/train_job.py",
68
+ flavor="rtx-pro-6000", timeout="1h",
69
+ secrets={"HF_TOKEN": TOKEN},
70
+ script_args=["--dataset", "<user>/<lora_name>-dataset",
71
+ "--lora-name", "<lora_name>",
72
+ "--trigger", "<trigger>", "--concept", "style"],
73
+ )
74
+ print(job.url)
75
+ ```
76
+
77
+ Useful flags: `--steps` (1000), `--rank` (32), `--resolution` (1024), `--learning-rate` (3e-4),
78
+ `--quantization {none,fp8,4bit}`, `--no-gallery`, `--num-gallery` (3). Run with `--help` for all.
79
+
80
+ ## Result
81
+ ~40 min on `rtx-pro-6000` (1000 steps, regional torch.compile). The LoRA is pushed to
82
+ `<user>/<lora_name>` with a preview gallery + README. Use it:
83
+
84
+ ```python
85
+ import torch
86
+ from diffusers import Krea2Pipeline
87
+ pipe = Krea2Pipeline.from_pretrained("krea/Krea-2-Turbo", torch_dtype=torch.bfloat16).to("cuda")
88
+ pipe.load_lora_weights("<user>/<lora_name>")
89
+ image = pipe("<trigger>, a fox in a snowy forest", num_inference_steps=8, guidance_scale=0.0).images[0]
90
+ image.save("out.png")
91
+ ```