Base
Core extraction function and supporting types for LLM-based structured extraction.
Classes:
| Name | Description |
|---|---|
SingleExtractionResult |
Result container for a single extraction call. |
TextOffsetValueError |
Raised when character offset arguments are invalid. |
Functions:
| Name | Description |
|---|---|
extract_from_text |
Extract structured output from text using an LLM. |
extract_from_text_lenient |
Wrapper around |
build_chat_messages |
Build a list of chat messages from system/user templates,
inserting document text and schema description where needed, using |
build_chat_message |
Build a single chat message from a template. |
strip_metadata |
Strip metadata wrappers from a JSON-parsed result. |
augment_metadata |
Recursively augment metadata wrapper dicts.
Uses just |
augment_metadata_node_with_evidence |
Augment a single metadata wrapper dict with evidence location info. |
add_response_content_callback |
Postprocessing callback to add response content to output. |
add_reasoning_content_callback |
Postprocessing callback to add reasoning content to output. |
add_structured_callback |
Postprocessing callback to parse and validate structured output. |
augment_and_strip_metadata_from_structured_callback |
Postprocessing callback to augment metadata and strip it from structured output. |
exception2error_msg |
Format an exception into short and long error message strings. |
check_utf8_encodable |
Recursively check that all nested strings can be encoded to UTF-8. |
TextOffsetValueError
Bases: ValueError
Raised when text offset is invalid.
SingleExtractionResult(character_start, character_end, response_content=None, structured=None, structured_with_metadata=None, reasoning_content=None, messages=None, messages_formatted=None, errors=list(), errors_long=list())
dataclass
Bases: FieldDict
Stores one extraction result
Attributes:
| Name | Type | Description |
|---|---|---|
character_start |
int
|
Start index of the chunk of text that has been processed |
character_end |
int
|
End index of the chunk of text that has been processed (exclusive) |
response_content |
str | None
|
Response formatted as json |
structured |
dict[str, Any] | list[Any] | None
|
Parsed version of response_content. Possibly validated against a schema. - Internally may come with metadata. |
structured_with_metadata |
dict[str, Any] | list[Any] | None
|
Parsed version of response_content, with enriched metadata. |
reasoning_content |
str | None
|
The LLMs reasoning output, as parsed by the api endpoint.
(Usually output between |
messages |
dict[str, str | None] | None
|
prompt messages with placeholders for input text and schema description
|
messages_formatted |
dict[str, str] | None
|
prompt messages with inserted input text and schema description
|
errors |
list[str]
|
list of strings "error_name: error_message" |
errors_long |
list[str]
|
list of strings "error_name: error_message_and_traceback" |
exception2error_msg(e)
Return short and long (including traceback) error messages for an exception.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
e
|
BaseException
|
Python exception object |
required |
Returns:
| Type | Description |
|---|---|
str
|
Tuple containing a short and a long version of the exception. |
str
|
The short version has the exception name and message, |
tuple[str, str]
|
whilst the long version also comes with the entire traceback. |
Source code in src/kibad_llm/extractors/base.py
94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 | |
check_utf8_encodable(data)
Recursively check that every string in data can be encoded to UTF-8.
Walks nested dicts/lists/tuples and encodes every string leaf. A lone (unpaired)
surrogate is valid in a Python str but cannot be encoded to UTF-8, which is
the reason datasets/pyarrow crash with an unhandled UnicodeEncodeError
when writing out a batch that contains one. Encoding here, right when the
offending string is produced, lets us attribute the failure to a single
document instead of losing an entire batch's results.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Any
|
Arbitrary nested data (e.g. a |
required |
Raises:
| Type | Description |
|---|---|
UnicodeEncodeError
|
If any string in |
Source code in src/kibad_llm/extractors/base.py
113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 | |
build_chat_message(message, role, document=None, document_placeholder='document', schema=None, schema_description_kwargs=None, schema_description_placeholder='schema_description')
Build a single chat message by inserting text and schema description if respective placeholders are present in the message template.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
message
|
str
|
The message template. |
required |
role
|
MessageRole
|
The role of the message (e.g., system, user). |
required |
document
|
str | None
|
The document text to process. |
None
|
document_placeholder
|
str
|
The placeholder in the message template for the input text. If the placeholder is present in the message template, it will be replaced with the input text. |
'document'
|
schema
|
dict[str, Any] | None
|
Optional JSON schema for structured output. |
None
|
schema_description_kwargs
|
dict[str, Any] | None
|
Optional kwargs for build_schema_description when generating the schema description. |
None
|
schema_description_placeholder
|
str
|
The placeholder in the message template for the schema description. If the placeholder is present in the message template, the schema must be provided and the description will be generated and inserted. |
'schema_description'
|
Returns:
| Type | Description |
|---|---|
SimpleChatMessage
|
A tuple of ChatMessage and a metadata dictionary indicating whether schema description |
dict[str, bool]
|
and text were inserted. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If a schema is required but not supplied. |
ValueError
|
If a document is required but not supplied. |
Source code in src/kibad_llm/extractors/base.py
148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 | |
build_chat_messages(system_message=None, user_message=None, schema_description_placeholder='schema_description', document_placeholder='document', schema=None, history=None, return_messages=False, return_messages_formatted=False, truncate_user_message_formatted=300, _out=None, **build_messages_kwargs)
Build chat messages for extraction. The document text and schema description may be inserted into the message templates, depending on the presence of the respective placeholders.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
system_message
|
str | None
|
The system message template. |
None
|
user_message
|
str | None
|
The user message template. |
None
|
schema
|
dict[str, Any] | None
|
Optional JSON schema for structured output. |
None
|
schema_description_placeholder
|
str
|
The placeholder in the message templates for the schema description. If the placeholder is present in the message templates, the schema must be provided and the description will be generated and inserted. |
'schema_description'
|
document_placeholder
|
str
|
The placeholder in the message templates for the input text. If the placeholder is present in the message templates, it will be replaced with the input text. |
'document'
|
history
|
list[SimpleChatMessage] | None
|
Optional list of ChatMessage objects representing the conversation history. |
None
|
return_messages
|
bool
|
Whether to return the used prompt messages, but without input text and schema description. |
False
|
return_messages_formatted
|
bool
|
Whether to return the used prompt messages formatted with input text and schema description. |
False
|
truncate_user_message_formatted
|
int | None
|
If return_messages_formatted is True, truncate the user message content to this many characters (to avoid huge outputs). Set to None to disable truncation. |
300
|
_out
|
SingleExtractionResult | None
|
Optional output |
None
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**build_messages_kwargs |
Any
|
Additional keyword arguments for build_chat_message. |
Returns:
| Type | Description |
|---|---|
list[SimpleChatMessage]
|
A list of ChatMessage objects. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If neither a system_message nor user_message are provided. |
ValueError
|
If no document placeholder is supplied and history is false. (no input text would be inserted) |
Warns:
| Type | Description |
|---|---|
UserWarning
|
If a schema description is supplied, but there is no schema description placeholder. |
Source code in src/kibad_llm/extractors/base.py
214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 | |
strip_metadata(data, *, content_key)
Strip metadata wrappers from a JSON-parsed result produced by wrap_terminals_with_metadata.
The wrapped output encodes terminal values as objects like
{"<content_key>": <value>, "evidence_anchor": "...", ...}
This function walks the parsed JSON (dict/list/scalars) and removes such wrappers by
replacing the wrapper dict with its <content_key> value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Any
|
Parsed JSON with metadata wrappers. |
required |
content_key
|
str
|
Key whose metadata wrappers to remove. - must be passed by keyword |
required |
Returns:
| Type | Description |
|---|---|
Any
|
Parsed JSON without metadata wrappers around content_key. |
Wrapper detection (heuristic):
- a dict is treated as a wrapper if it has content_key AND at least one additional key.
(We avoid unwrapping objects that only have {"<content_key>": ...}.)
Notes
- This function does not validate that "other keys" are truly metadata. If your original
extraction schema contains real objects that also have a
content_keyfield and other fields, they may be unwrapped unintentionally. If that’s a concern, use a more uniquecontent_key(e.g. "__content") in the schema wrapping step. - The input is not mutated; a transformed copy is returned.
Source code in src/kibad_llm/extractors/base.py
336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 | |
augment_metadata_node_with_evidence(node, text, token_spans, *, anchor_key='evidence_anchor', num_matches_key='evidence_num_matches', start_key='first_evidence_start', end_key='first_evidence_end', snippet_key='first_evidence_snippet', snippet_margin=10, character_offset=0)
Augment a single metadata wrapper dict with evidence location information.
Given a wrapper object like
{"content": ..., "evidence_anchor": "...", ...}
this function searches text for the anchor (via _find_anchor_match_spans). If at least
one match is found, it adds:
- num_matches_key: number of matches
- start_key / end_key: character offsets of the first match
- snippet_key: a substring of text spanning snippet_margin tokens around the match
(whitespace preserved)
If no anchor is present (or no matches exist), the wrapper is returned unchanged except for
num_matches_key (only added when an anchor is a non-empty string).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node
|
Mapping[str, Any]
|
The metadata wrapper dict to augment. |
required |
text
|
str
|
The original text to search for evidence anchors. |
required |
token_spans
|
list[tuple[int, int]]
|
Precomputed list of (start_offset, end_offset) tuples for each token in |
required |
anchor_key
|
str
|
The key in wrapper dicts that holds the evidence anchor text. |
'evidence_anchor'
|
num_matches_key
|
str
|
The key to add for the number of matches of the anchor in the text. |
'evidence_num_matches'
|
start_key
|
str
|
The key to add for the start character offset of the anchor. |
'first_evidence_start'
|
end_key
|
str
|
The key to add for the end character offset of the anchor. |
'first_evidence_end'
|
snippet_key
|
str
|
The key to add for the evidence snippet text. |
'first_evidence_snippet'
|
snippet_margin
|
int
|
Number of tokens to include before and after the anchor span in the snippet. |
10
|
character_offset
|
int
|
An optional offset to add to the start/end character positions (useful if the text is a chunk of a larger document). |
0
|
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
The augmented metadata wrapper dict with evidence metadata added where applicable. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the snippet_margin is negative. |
Source code in src/kibad_llm/extractors/base.py
503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 | |
augment_metadata(data, *, text, content_key, **kwargs)
Recursively augment all metadata wrapper dicts in a JSON-parsed result.
Currently, this does the following node processing: - evidence anchor handling using augment_metadata_node_with_evidence
Traversal
- walks
datathrough nested dicts/lists - detects wrapper dicts via
_is_wrapper_dict(..., content_key=...) - for each wrapper dict, calls metadata node augmentation methods
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
Any
|
JSON-parsed result with metadata |
required |
text
|
str
|
Original input text to extract further information from - must be passed by keyword |
required |
content_key
|
str
|
Key that must be in a mapping for it to be a wrapper_dict - must be passed by keyword |
required |
Other Parameters:
| Name | Type | Description |
|---|---|---|
evidence_ |
Any
|
|
Keyword Args Notes
- kwargs are namespaced by prefix.
Example: evidence_snippet_margin=10 sets
snippet_margin=10for evidence augmentation.
Returns:
| Type | Description |
|---|---|
Any
|
The data with the augmented metadata. |
Any
|
The returned structure mirrors the input but includes added metadata fields where applicable. |
Raises:
| Type | Description |
|---|---|
ValueError
|
unknown kwargs raise ValueError (fail fast). |
Source code in src/kibad_llm/extractors/base.py
577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 | |
add_response_content_callback(out, response, *, llm)
Add response_content to output dictionary.
Modifies out in place; does not return a value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
out
|
SingleExtractionResult
|
SingleExtractionResult to store the response_content in. |
required |
response
|
ChatResponse
|
ChatResponse that contains the response_content. |
required |
llm
|
LLM
|
LLM object that extracts the response_content from response. - must be passed by keyword |
required |
Source code in src/kibad_llm/extractors/base.py
651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 | |
add_reasoning_content_callback(out, response, *, llm)
Add reasoning_content to output dictionary.
Modifies out in place; does not return a value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
out
|
SingleExtractionResult
|
SingleExtractionResult to store the reasoning_content in. |
required |
response
|
ChatResponse
|
ChatResponse that contains the reasoning_content. |
required |
llm
|
LLM
|
LLM object that extracts the reasoning_content from response. - must be passed by keyword |
required |
Source code in src/kibad_llm/extractors/base.py
669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 | |
add_structured_callback(out, response, *, schema, validate_with_schema)
Add structured output to output dictionary based on response content.
Modifies out in place; does not return a value.
This structured output may or may not contain metadata to some extent.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
out
|
SingleExtractionResult
|
SingleExtractionResult to store the structured output in. |
required |
response
|
ChatResponse
|
Raw ChatResponse from the LLM call. - This exists for compatibility reasons and is not used here. |
required |
schema
|
dict[str, Any] | None
|
Schema to validate the response against. - must be passed by keyword |
required |
validate_with_schema
|
bool
|
Whether to validate response against schema. - must be passed by keyword |
required |
Source code in src/kibad_llm/extractors/base.py
687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 | |
augment_and_strip_metadata_from_structured_callback(out, response, *, schema, original_schema, text, validate_with_schema, augment_metadata_kwargs=None)
Augment metadata in structured output and save it as structured_with_metadata.
Then, strip metadata and save the cleaned version back to structured.
Modifies out in place; does not return a value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
out
|
SingleExtractionResult
|
An extraction result with |
required |
response
|
ChatResponse
|
Placeholder, because it gets passed to all functions in |
required |
schema
|
dict[str, Any] | None
|
schema used in |
required |
original_schema
|
dict[str, Any] | None
|
original schema that was passed to |
required |
text
|
str
|
Original input text for information extraction - must be passed by keyword |
required |
validate_with_schema
|
bool
|
Whether to validate |
required |
augment_metadata_kwargs
|
dict[str, Any] | None
|
Refer to |
None
|
Warning
requires out to have structured, so run ..add_structured_callback first
Source code in src/kibad_llm/extractors/base.py
717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 | |
extract_from_text(text, text_id, prompt_template, schema=None, use_guided_decoding=True, validate_with_schema=True, llm=None, request_parameters=None, return_reasoning=False, adjust_schema_for_evidence_detection=False, adjust_schema_description_for_evidence_detection=False, evidence_anchor_description='Verbatim excerpt from the source text supporting the extracted content.', wrapped_content_description=None, response_has_metadata=False, augment_metadata_kwargs=None, character_start=0, character_end=None, user_message=None, system_message=None, schema_description_placeholder=None, document_placeholder=None, **build_messages_kwargs)
Extract structured information from text using an LLM.
Given a chat llm, composes system and user messages, and invokes the model. When a schema is provided, it is used to enforce guided decoding. The output is parsed as JSON and validated against the schema if provided.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The document text to process. |
required |
text_id
|
str
|
Document text identifier for logging. |
required |
prompt_template
|
dict[str, str | None]
|
A dictionary with at least one of 'system_message' and 'user_message' templates (or both). |
required |
schema
|
dict[str, Any] | None
|
Optional JSON schema for structured output. |
None
|
use_guided_decoding
|
bool
|
Whether to use guided decoding. |
True
|
validate_with_schema
|
bool
|
Whether to validate the output against the provided schema. IMPORTANT: Disabling validation may lead to invalid structured outputs and, thus, may break result serialization (since we use .map() and .to_json() from datasets). |
True
|
llm
|
LLM | None
|
The LLM model to use. Must be a chat model (i.e. is_chat_model=True) and support extra_body parameters for guided decoding if schema is provided. If None, no LLM call is made. |
None
|
request_parameters
|
dict[str, Any] | None
|
Additional parameters to pass to the LLM chat call. |
None
|
return_reasoning
|
bool
|
Whether to return the reasoning done by the model. |
False
|
adjust_schema_for_evidence_detection
|
bool
|
Whether to adjust the schema to wrap terminal values
with metadata. If True, the schema is modified so that each terminal value is replaced
with an object containing the original value under the key |
False
|
adjust_schema_description_for_evidence_detection
|
bool
|
Whether to adjust the schema description when detect_evidence is True. If True, the schema description will mention that each value is accompanied by an evidence_anchor that is a "verbatim excerpt from the source text supporting the extracted content" (see METADATA_SCHEMA_WITH_EVIDENCE_SHORTHAND). Has only an effect if adjust_schema_for_detect_evidence is also True. |
False
|
evidence_anchor_description
|
str
|
Description for the evidence anchor field. |
'Verbatim excerpt from the source text supporting the extracted content.'
|
wrapped_content_description
|
str | None
|
Optional description for the content field in the metadata wrapper. |
None
|
response_has_metadata
|
bool
|
If True, the output is expected to have each leaf value wrapped in
an object with |
False
|
augment_metadata_kwargs
|
dict[str, Any] | None
|
Additional keyword arguments for augment_metadata. |
None
|
character_start
|
int
|
Optional character offsets to specify a substring of |
0
|
character_end
|
int | None
|
Optional character offsets to specify a substring of |
None
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**build_messages_kwargs |
Any
|
Additional keyword arguments for build_chat_messages. |
Returns:
| Type | Description |
|---|---|
SingleExtractionResult
|
A SingleExtractionResult object with the extraction result. |
Warns:
| Type | Description |
|---|---|
UserWarning
|
When there is no LLM provided, the call to it is skipped. |
Raises:
| Type | Description |
|---|---|
DeprecationWarning
|
If deprecated args are used, the extraction exits early. |
TextOffsetValueError
|
If character_start and/or character_end are set erroneously. |
ValueError
|
If schema is None and use_guided_decoding or adjust_schema_for_evidence_detection is True. |
Source code in src/kibad_llm/extractors/base.py
779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 | |
extract_from_text_lenient(text, text_id, character_start=0, character_end=None, **kwargs)
Wrapper around extract_from_text that catches all exceptions.
This is useful when processing multiple documents and we want to continue processing even if one document fails.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The text to process. |
required |
text_id
|
str
|
Text identifier for logging. |
required |
character_start
|
int
|
Optional character offsets to specify a substring of |
0
|
character_end
|
int | None
|
Optional character offsets to specify a substring of |
None
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
**kwargs |
Any
|
Keyword arguments for |
Returns:
| Type | Description |
|---|---|
SingleExtractionResult
|
A SingleExtractionResult object with the extraction result or error message. |
Source code in src/kibad_llm/extractors/base.py
1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 | |