Multi pass
MultiPassExtractorWithChunking chunking based extraction in multiple passes.
Classes:
| Name | Description |
|---|---|
MultiPassExtractorWithChunking |
Combines |
MultiPassExtractorWithChunking(overrides, aggregator, chunking_aggregator, return_as_list=None, tokenizer=None, max_char_buffer=20000, **kwargs)
Extractor that executes one extraction pass for each entry in a list (or dict) of parameter overrides. In each pass, the input text is chunked and the base extraction function is called on each chunk. The results of the chunks are aggregated per pass, and then the results of all passes are aggregated into a single result.
Attributes:
| Name | Type | Description |
|---|---|---|
overrides |
A list of dictionaries containing parameter overrides for each extraction pass. The number of entries defines the number of passes. Can also be a dictionary in the format {"pass id" -> "override parameters"} to improve config readability (the pass id is not used for anything else). |
|
aggregator |
Aggregator function to combine the results of the passes (outer loop). |
|
chunking_aggregator |
Aggregator function to combine the results of the chunks within a single pass (inner loop). |
|
return_as_list |
List of field names to return as lists of all extracted values. Length will be the number of extraction passes (override entries x chunks). |
|
tokenizer |
Tokenizer to use for chunking. |
|
max_char_buffer |
Max chunk size in characters. |
|
default_kwargs |
Additional keyword arguments passed to the base extraction function. |
Warning
If a Token that is greater than max_char_buffer is encountered, it becomes its own chunk. This edge case can produce chunks that are larger than max_char_buffer would allow.
See UnionExtractor as well as
ChunkingExtractor for accepted parameters
and details about the aggregation logic.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
overrides
|
list[dict] | dict[str, dict]
|
A list of dictionaries containing parameter overrides for each extraction pass. Can also be a dictionary in the format {"pass id" -> "override parameters"} to improve config readability (the pass id is not used for anything else). |
required |
aggregator
|
Aggregator
|
Aggregator function to use across passes (outer loop). |
required |
chunking_aggregator
|
Aggregator
|
Aggregator function to use across chunks (inner loop). |
required |
return_as_list
|
list[str] | None
|
List of field names to return as lists of all extracted values |
None
|
tokenizer
|
Tokenizer | None
|
Tokenizer to use for chunking. |
None
|
max_char_buffer
|
int
|
Max chunk size in characters. |
20000
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
* |
Additional keyword arguments passed to the base extraction function. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If no overrides are supplied, there can't be an extraction. |
Source code in src/kibad_llm/extractors/multi_pass.py
46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 | |
__call__(*args, **kwargs)
Process singular text in chunks with multiple passes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
*args
|
Any
|
Positional form of |
()
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
text |
str
|
Input document to process. |
text_id |
str
|
Id of input document. |
* |
Any
|
Refer to |
Returns:
| Name | Type | Description |
|---|---|---|
dict[str, Any]
|
Dict with the key |
|
dict[str, Any]
|
Additionally, there can be lists for fields at the keys |
|
dict[str, Any]
|
one entry per extraction call, i.e. one per (pass, chunk) pair in pass-major |
|
order |
dict[str, Any]
|
all chunks of the first pass, then all chunks of the second, and so on. |
dict[str, Any]
|
This flattens the per-call layout of |
|
dict[str, Any]
|
|
|
dict[str, Any]
|
|
Source code in src/kibad_llm/extractors/multi_pass.py
86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 | |