Skip to content

Multi pass

MultiPassExtractorWithChunking chunking based extraction in multiple passes.

Classes:

Name Description
MultiPassExtractorWithChunking

Combines UnionExtractor and ChunkingExtractor by running one pass per override and, within each pass, the ChunkingExtractor loop over the chunks of the text.

MultiPassExtractorWithChunking(overrides, aggregator, chunking_aggregator, return_as_list=None, tokenizer=None, max_char_buffer=20000, **kwargs)

Extractor that executes one extraction pass for each entry in a list (or dict) of parameter overrides. In each pass, the input text is chunked and the base extraction function is called on each chunk. The results of the chunks are aggregated per pass, and then the results of all passes are aggregated into a single result.

Attributes:

Name Type Description
overrides

A list of dictionaries containing parameter overrides for each extraction pass. The number of entries defines the number of passes. Can also be a dictionary in the format {"pass id" -> "override parameters"} to improve config readability (the pass id is not used for anything else).

aggregator

Aggregator function to combine the results of the passes (outer loop).

chunking_aggregator

Aggregator function to combine the results of the chunks within a single pass (inner loop).

return_as_list

List of field names to return as lists of all extracted values. Length will be the number of extraction passes (override entries x chunks).

tokenizer

Tokenizer to use for chunking.

max_char_buffer

Max chunk size in characters.

default_kwargs

Additional keyword arguments passed to the base extraction function.

Warning

If a Token that is greater than max_char_buffer is encountered, it becomes its own chunk. This edge case can produce chunks that are larger than max_char_buffer would allow.

See UnionExtractor as well as ChunkingExtractor for accepted parameters and details about the aggregation logic.

Parameters:

Name Type Description Default
overrides list[dict] | dict[str, dict]

A list of dictionaries containing parameter overrides for each extraction pass. Can also be a dictionary in the format {"pass id" -> "override parameters"} to improve config readability (the pass id is not used for anything else).

required
aggregator Aggregator

Aggregator function to use across passes (outer loop).

required
chunking_aggregator Aggregator

Aggregator function to use across chunks (inner loop).

required
return_as_list list[str] | None

List of field names to return as lists of all extracted values

None
tokenizer Tokenizer | None

Tokenizer to use for chunking.

None
max_char_buffer int

Max chunk size in characters.

20000

Other Parameters:

Name Type Description
*

Additional keyword arguments passed to the base extraction function.

Raises:

Type Description
ValueError

If no overrides are supplied, there can't be an extraction.

Source code in src/kibad_llm/extractors/multi_pass.py
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
def __init__(
    self,
    overrides: list[dict] | dict[str, dict],
    aggregator: Aggregator,
    chunking_aggregator: Aggregator,
    return_as_list: list[str] | None = None,
    tokenizer: tokenizer_lib.Tokenizer | None = None,
    max_char_buffer: int = 20000,
    **kwargs,
):
    """Assign args to attributes with safety checks and conversions

    Args:
        overrides: A list of dictionaries containing parameter overrides for each extraction pass.
            Can also be a dictionary in the format {"pass id" -> "override parameters"} to improve
            config readability (the pass id is not used for anything else).
        aggregator: Aggregator function to use across passes (outer loop).
        chunking_aggregator: Aggregator function to use across chunks (inner loop).
        return_as_list: List of field names to return as lists of all extracted values
        tokenizer: Tokenizer to use for chunking.
        max_char_buffer: Max chunk size in characters.

    Keyword Args:
        *: Additional keyword arguments passed to the base extraction function.

    Raises:
        ValueError: If no overrides are supplied, there can't be an extraction.
    """
    if len(overrides) < 1:
        raise ValueError("overrides must contain at least one set of parameters")
    if isinstance(overrides, list):
        overrides = {str(i): override for i, override in enumerate(overrides)}
    self.overrides = overrides
    self.aggregator = aggregator
    self.chunking_aggregator = chunking_aggregator
    self.return_as_list = return_as_list or []
    self.default_kwargs = kwargs
    self.tokenizer = tokenizer
    self.max_char_buffer = max_char_buffer

__call__(*args, **kwargs)

Process singular text in chunks with multiple passes.

Parameters:

Name Type Description Default
*args Any

Positional form of text and text_id, in that order.

()

Other Parameters:

Name Type Description
text str

Input document to process.

text_id str

Id of input document.

* Any

Returns:

Name Type Description
dict[str, Any]

Dict with the key structured that holds the aggregated structured outputs.

dict[str, Any]

Additionally, there can be lists for fields at the keys "{field}_list". These hold

dict[str, Any]

one entry per extraction call, i.e. one per (pass, chunk) pair in pass-major

order dict[str, Any]

all chunks of the first pass, then all chunks of the second, and so on.

dict[str, Any]

This flattens the per-call layout of

dict[str, Any]

UnionExtractor on the outside and

dict[str, Any]

ChunkingExtractor on the inside.

Source code in src/kibad_llm/extractors/multi_pass.py
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
def __call__(self, *args, **kwargs) -> dict[str, Any]:
    """Process singular text in chunks with multiple passes.

    Args:
        *args (Any): Positional form of `text` and `text_id`, in that order.


    Keyword Args:
        text (str): Input document to process.
        text_id (str): Id of input document.
        * (Any): Refer to [`extract_from_text_lenient`][kibad_llm.extractors.base.extract_from_text_lenient]

    Returns:
        Dict with the key `structured` that holds the aggregated structured outputs.
        Additionally, there can be lists for fields at the keys `"{field}_list"`. These hold
        one entry per extraction call, i.e. one per (pass, chunk) pair in pass-major
        order: all chunks of the first pass, then all chunks of the second, and so on.
        This flattens the per-call layout of
        [`UnionExtractor`][kibad_llm.extractors.union.UnionExtractor] on the outside and
        [`ChunkingExtractor`][kibad_llm.extractors.chunking.ChunkingExtractor] on the inside.
    """

    # extract text and id in most compatible way
    text = kwargs.pop("text", None)
    if text is None:
        text = args[0]

    text_id = kwargs.pop("text_id", None)
    if text_id is None:
        text_id = args[-1]

    combined_kwargs = {**self.default_kwargs, **kwargs}

    # chunk the input text first so that we don't have to do it for every pass
    chunks = _document_chunk_iterator(
        document=text,
        max_char_buffer=self.max_char_buffer,
        tokenizer=self.tokenizer,
    )

    # the raw results of every single extraction call (one per chunk and pass)
    # to build the "{field}_list" entries from
    all_results = []

    # aggregated structured output per pass
    all_passes_results = []
    for override_name, override_params in self.overrides.items():
        current_pass_results = []
        current_pass_kwargs = {**combined_kwargs, **override_params}
        for i, chunk in enumerate(chunks):
            # for each override, run the ChunkingExtractor loop over all chunks of the text
            current_chunk_result = extract_from_text_lenient(
                text=text,
                text_id=f"{text_id}_{override_name}_chunk_{i}",
                **current_pass_kwargs,
                # This may raise an error if character_start or character_end is already provided via kwargs,
                # but we want to be strict about not allowing that since it would interfere with the chunking logic.
                character_start=chunk.char_interval.start_pos or 0,
                character_end=chunk.char_interval.end_pos,
            )

            current_pass_results.append(current_chunk_result)

        all_results.extend(current_pass_results)

        # aggregate the results of the current override's extraction passes over all chunks
        current_pass_structured_outputs = [
            v.get("structured", None) for v in current_pass_results
        ]
        all_passes_results.append(self.chunking_aggregator(current_pass_structured_outputs))

    # aggregate the previously aggregated results, but now across the overrides, to get a single result.
    aggregated_structured = self.aggregator(all_passes_results)

    result: dict[str, Any] = {
        "structured": aggregated_structured,
    }
    for field in self.return_as_list:
        result[f"{field}_list"] = [v.get(field, None) for v in all_results]
    return result