Chunking
ChunkingExtractor for running
extraction over bounded chunks of a document.
Classes:
| Name | Description |
|---|---|
ChunkingExtractor |
Splits a document into chunks and aggregates per-chunk extraction results. |
ChunkingExtractor(aggregator, return_as_list=None, tokenizer=None, max_char_buffer=20000, verbose=False, **kwargs)
Extractor that chunks extraction and aggregates results per key. This extractor calls the base extraction function multiple times (for each chunk in the document) on the same input text, passing no previous context to each subsequent call.
Pass llm=None with verbose=True to get the number of chunks per document without inference.
Attributes:
| Name | Type | Description |
|---|---|---|
aggregator |
Method to aggregate the llm output for the individual chunks before returning |
|
return_as_list |
List of field names to return as lists of all extracted values |
|
tokenizer |
tokenizer to use for chunking |
|
max_char_buffer |
Max chunk size in characters |
|
verbose |
Adds verbose logging |
|
default_kwargs |
Additional keyword arguments passed to the base extraction function. |
Warning
If a Token that is greater than max_char_buffer is encountered, it becomes its own chunk. This edge case can produce chunks that are larger than max_char_buffer would allow.
Source code in src/kibad_llm/extractors/chunking.py
66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 | |
__call__(*args, **kwargs)
Processes a text in chunks.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
*args
|
Any
|
Positional form of |
()
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
text |
str
|
Input document to process. |
text_id |
str
|
Id of input document. |
* |
Any
|
Refer to |
Returns:
| Type | Description |
|---|---|
dict[str, Any]
|
Dict with the key |
dict[str, Any]
|
Additionally, there can be lists for fields at the keys |
Source code in src/kibad_llm/extractors/chunking.py
82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | |