Quickstart
A cheat sheet for getting kibad-llm running and understanding how it fits together. This project uses LLMs
to extract structured information from scientific literature PDFs, supporting the
Literaturdatenbank Faktencheck Artenvielfalt.
For anything not covered here, see Where to go deeper below.
Setup
Requires uv (install guide).
git clone https://github.com/DFKI-NLP/kibad-llm
cd kibad-llm
uv sync --group cicd # installs the project plus dev/CI tooling (lint, test, docs)
cp .env.example .env # then fill in the variables you need, see below
.env variables (all optional, only needed for the features that use them):
OPENAI_API_KEY— required to run extraction with OpenAI-hosted models (e.g.gpt_5). Create a key at platform.openai.com/api-keys.HF_TOKEN— required for access-restricted Hugging Face models (e.g.gemma3_27b), and for the in-process vLLM backend. Create a token at huggingface.co/settings/tokens.VLLM_DOWNLOAD_DIR— optional, where in-process vLLM caches downloaded model weights.DB_USER/DB_PASSWORD— only needed for the Faktencheck Postgres database, see podman/faktencheck-db/README.md.
[!TIP] If you're new to
uv: wherever you'd normally writepython ..., writeuv run python ...(or, for the project's own CLI entry points,uv run -m kibad_llm....) — no need to manually activate a virtualenv.
Getting data in
- PDFs from Zotero:
uv run -m kibad_llm.data_integration.zotero_downloaddownloads open-access papers found via Semantic Scholar from an exported Zotero-group CSV (see data/external/zotero). Details: USAGE.md § PDF Download. - Faktencheck reference data: the ground-truth database lives in Postgres. Start it with Podman (see
podman/faktencheck-db/README.md), then convert it to JSON with
uv run -m kibad_llm.data_integration.db_converter. Scientific names in the converted data can be normalized against GBIF viauv run -m kibad_llm.normalization.cli gbif .... Details: USAGE.md § Faktencheck Postgres to Json Conversion. - What's already available: data/readme.md documents the existing PDF sets and reference/ground-truth files, so check there before re-downloading or re-converting anything.
Running the extraction pipeline
Prerequisite: host an LLM
Prediction needs a running LLM backend, chosen per extractor config under configs/extractor/llm/:
- OpenAI-hosted (e.g.
gpt_5) — just needsOPENAI_API_KEY, no separate hosting step. - External vLLM server (
*_in_process.yaml's sibling, an OpenAI-compatible endpoint you start yourself) — see models/README.md andrun_with_llm.sh. - In-process vLLM (
*_in_process.yaml, model loaded inside the same process; used on the DFKI cluster viarun_in_process.sh) — needsHF_TOKENfor gated models.
Follow models/README.md § Quickstart or § All-in-one run script to get one running.
Run a prediction
uv run -m kibad_llm.predict pdf_directory=path/to/pdf/files
Converts every PDF in pdf_directory to markdown, runs the configured extractor over it, and writes a JSONL
predictions file. Options live in configs/predict.yaml; pdf_reader_num_proc and
extractor_num_proc control parallelism (keep them modest on shared/personal machines, large on compute
nodes). For a reproducible setup, use a named config instead of ad hoc overrides, e.g. the organism-trends
extraction:
uv run -m kibad_llm.predict pdf_directory=path/to/pdf/files experiment/predict=organism_trends
See configs/experiment/predict for available experiment configs and USAGE.md § Inference for the full picture.
Run an evaluation
uv run -m kibad_llm.evaluate dataset.predictions.file=path/to/predictions.jsonl
Scores predictions against reference data. Defaults to dataset=faktencheck and metric=f1_micro
(micro-averaged precision/recall/F1 over all fields); swap either via dataset=<name> /
metric=<name> — see configs/dataset and configs/metric. As with
prediction, prefer a named experiment/evaluate=<config> for anything you'll want to reproduce (see
configs/experiment/evaluate). Full details, including the
confusion_matrix metric's per-field requirement:
USAGE.md § Evaluation.
Multirun and A/B testing
Both entry points support Hydra multirun:
pass comma-separated values for one or more parameters plus --multirun (-m), and Hydra runs every
combination — e.g. extractor=simple_with_schema,simple --multirun to compare guided vs. unguided decoding, or
seed=42,1337,7331 --multirun for repeated runs. Add
+hydra.callbacks.save_job_return.multirun_markdown_group_by=<column(s)> to evaluate to get aggregated
mean/std across runs. See USAGE.md § Multirun for worked examples.
Inspect results
Runs write to logs/<name>/... and predictions/<name>/... locally (git-ignored). Each run produces a
job_return_value.json/.md with output paths or metric scores; multiruns get a combined summary. Finished
experiments are committed to the separate kibad-llm-results
repository instead, which you can clone to data/results — see Repo conventions below.
Browse and compare runs (local, from the repo, or from a GitHub
URL) in the build-free evaluation dashboard.
How it fits together
Everything is wired through Hydra config groups under configs/ via _target_, so adding a new
LLM, extractor, metric, or dataset means adding both a Python class/function and a matching YAML.
- Extractors (src/kibad_llm/extractors/) all share the contract
(text, file_name) -> dictand compose: a single LLM call at the core (builds a prompt from a schema derived from the pydantic models in src/kibad_llm/schema/types.py, optionally with guided decoding), wrapped byChunkingExtractor(splits long documents),UnionExtractor/ConditionalUnionExtractor(multiple passes, merged or chained), andRepeatingExtractor(majority vote). - LLM backends (src/kibad_llm/llms/) share one interface over the three hosting
options described above (
openai.py,openai_like_vllm.py,vllm_in_process.py). - Evaluation (src/kibad_llm/evaluate.py) pairs a
dataset(predictions + references, matched on file name / record id, see src/kibad_llm/dataset/) with ametricimplementingreset/update/compute(see src/kibad_llm/metrics/). - Data integration (src/kibad_llm/data_integration/) holds the standalone Zotero/Postgres/GBIF/Nextcloud scripts mentioned above — these aren't part of the extraction pipeline itself.
Testing and before a PR
just pr # prek (lint/format/docs) + the full test suite — run this before claiming CI-readiness
just prek # just the lint/format/docs checks
just pytest # just the Python tests
just node-test # eval-dashboard JS logic tests (needs Node)
just prop # serve the docs locally
just -l lists all recipes. Link checking (lychee) runs as part of just prek and therefore just pr; the
standalone just lychee recipe is only a convenience wrapper. A single test: uv run pytest tests/unit/extractors/test_base.py::test_extract_from_text. Tests hitting a real LLM must
be marked slow (excluded from the default run) — prefer the llm_chat_replay fixture instead. If you touch a
test that uses recorded fixtures, regenerate only that test's fixtures:
WRITE_FIXTURE_DATA=1 uv run --group cicd pytest tests/integration/test_extractors.py tests/integration/test_predict.py
uv run --group cicd python tests/fixtures/map_llm_chat_usage.py # mandatory afterwards: flags now-unused fixtures
Never set WRITE_FIXTURE_DATA/WRITE_LLM_CHAT_FIXTURE_DATA for a full test-suite run — only for the tests you
intentionally changed.
Repo conventions
- Docstrings are mandatory on every file, class, function, and method — Google-style, CommonMark only (no Sphinx/reST). See CONTRIBUTING-CODE.md § Documentation.
- Tests mirror source layout:
tests/unit/mirrorssrc/kibad_llm/,tests/integration/mirrorsconfigs/and prefers real Hydra configs over mocks. - Branch naming:
feat/,fix/,hotfix/,docs/,experiment/prefixes, alphanumeric + hyphens only. Pushing tomainis prohibited; PRs are reviewed and squash-merged. - Experiment results live in a separate repository, kibad-llm-results
(committed logs and predictions). Clone it into the gitignored
data/results; branch and open PRs there as described in CONTRIBUTING-EXPERIMENTS.md. uv.lockis managed viauv add/uv lock, never hand-edited; explain dependency changes in the PR.- Windows:
uv sync --group cicdfails there (vllm→rayships nowin_amd64wheels). Run lint/test/docs commands on Linux/macOS, WSL, or the cluster. - For planning, naming, and documenting a full reproducible experiment (not just a one-off run), see CONTRIBUTING-EXPERIMENTS.md.
Where to go deeper
- USAGE.md — the full walkthrough: PDF download, DB conversion, prediction, evaluation, multirun and A/B testing, all with complete option lists.
- CONTRIBUTING.md — full directory map, PR workflow, docs-site rules, and the complete local-CI command set.
- CONTRIBUTING-CODE.md — coding principles, test layout, docstring/linking conventions, fixture regeneration, dependency changes.
- CONTRIBUTING-EXPERIMENTS.md — how to plan, name, run, and document a reproducible experiment.
- data/readme.md — description of the datasets and reference files available.
- dfki-nlp.github.io/kibad-llm — the rendered documentation site with all of the above, plus the auto-generated code reference.