Skip to content

Core

Library for breaking documents into chunks of sentences.

When a text-to-text model (e.g. a large language model with a fixed context size) can not accommodate a large document, this library can help us break the document into chunks of a required maximum length that we can perform inference on.

CharInterval(start_pos=None, end_pos=None) dataclass

Class for representing a character interval.

Attributes:

Name Type Description
start_pos int | None

The starting position of the interval (inclusive).

end_pos int | None

The ending position of the interval (exclusive).

TokenUtilError

Bases: BaseException

Error raised when token_util returns unexpected values.

TextChunk(token_interval, tokenized_text, document_text=None) dataclass

Stores a text chunk with attributes to the source document.

Attributes:

Name Type Description
token_interval TokenInterval

The token interval of the chunk in the source document.

tokenized_text TokenizedText

The source document in its tokenized form as TokenizedText obj.

document TokenizedText

The source document.

Properties: get_tokenized_text: TokenizedText of the current document. chunk_text: Text of the current chunk as string. sanitized_chunk_text Text of the current chunk as a sanitized string. -> _sanitize: Converts all whitespace characters in input text to a single space. char_interval: CharInterval of the chunk.

get_tokenized_text property

Gets the tokenized text from the source document.

chunk_text property

Gets the chunk text. Raises an error if document_text is not set.

sanitized_chunk_text property

Gets the sanitized chunk text.

char_interval property

Gets the character interval corresponding to the token interval.

Returns:

Type Description
CharInterval

data.CharInterval: The character interval for this chunk.

Raises:

Type Description
ValueError

If document_text is not set.

SentenceIterator(tokenized_text, curr_token_pos=0)

Iterate through sentences of a tokenized text.

Parameters:

Name Type Description Default
tokenized_text TokenizedText

Document to iterate through.

required
curr_token_pos int

Iterate through sentences from this token position.

0

Raises:

Type Description
IndexError

if curr_token_pos is not within the document.

Source code in src/kibad_llm/extractors/chunking_utils/core.py
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
def __init__(
    self,
    tokenized_text: tokenizer_lib.TokenizedText,
    curr_token_pos: int = 0,
):
    """Constructor.

    Args:
      tokenized_text: Document to iterate through.
      curr_token_pos: Iterate through sentences from this token position.

    Raises:
      IndexError: if curr_token_pos is not within the document.
    """
    self.tokenized_text = tokenized_text
    self.token_len = len(tokenized_text.tokens)
    if curr_token_pos < 0:
        raise IndexError(f"Current token position {curr_token_pos} can not be negative.")
    elif curr_token_pos > self.token_len:
        raise IndexError(
            f"Current token position {curr_token_pos} is past the length of the "
            f"document {self.token_len}."
        )
    self.curr_token_pos = curr_token_pos

__next__()

Returns next sentence's interval starting from current token position.

Returns:

Type Description
TokenInterval

Next sentence token interval starting from current token position.

Raises:

Type Description
StopIteration

If end of text is reached.

Source code in src/kibad_llm/extractors/chunking_utils/core.py
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
def __next__(self) -> tokenizer_lib.TokenInterval:
    """Returns next sentence's interval starting from current token position.

    Returns:
      Next sentence token interval starting from current token position.

    Raises:
      StopIteration: If end of text is reached.
    """
    assert self.curr_token_pos <= self.token_len
    if self.curr_token_pos == self.token_len:
        raise StopIteration
    # This locates the sentence which contains the current token position.
    sentence_range = tokenizer_lib.find_sentence_range(
        self.tokenized_text.text,
        self.tokenized_text.tokens,
        self.curr_token_pos,
    )
    assert sentence_range
    # Start the sentence from the current token position.
    # If we are in the middle of a sentence, we should start from there.
    sentence_range = create_token_interval(self.curr_token_pos, sentence_range.end_index)
    self.curr_token_pos = sentence_range.end_index
    return sentence_range

ChunkIterator(document, max_char_buffer, tokenizer_impl)

Iterate through chunks of a tokenized text.

Chunks may consist of sentences or sentence fragments that can fit into the maximum character buffer that we can run inference on.

Chunk cases:

A) If a sentence length exceeds the max char buffer, then it needs to be broken into chunks that can fit within the max char buffer. We do this in a way that maximizes the chunk length while respecting newlines (if present) and token boundaries. Consider this sentence from a poem by John Donne:

No man is an island,
Entire of itself,
Every man is a piece of the continent,
A part of the main.

With max_char_buffer=40, the chunks are: * "No man is an island,\nEntire of itself," len=38 * "Every man is a piece of the continent," len=38 * "A part of the main." len=19

B) If a single token exceeds the max char buffer, it comprises the whole chunk. Consider the sentence: "This is antidisestablishmentarianism." With max_char_buffer=20, the chunks are: * "This is" len=7 * "antidisestablishmentarianism" len=28 * "." len(1)

C) If multiple whole sentences can fit within the max char buffer, then they are used to form the chunk. Consider the sentences: "Roses are red. Violets are blue. Flowers are nice. And so are you." With max_char_buffer=60, the chunks are: * "Roses are red. Violets are blue. Flowers are nice." len=50 * "And so are you." len=15

Parameters:

Name Type Description Default
document str

Document to chunk. Can be either a string or a tokenized text.

required
max_char_buffer int

Size of buffer that we can run inference on.

required
tokenizer_impl Tokenizer

Tokenizer instance to use.

required
Source code in src/kibad_llm/extractors/chunking_utils/core.py
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
def __init__(
    self,
    document: str,
    max_char_buffer: int,
    tokenizer_impl: tokenizer_lib.Tokenizer,
):
    """Constructor.

    Args:
        document: Document to chunk. Can be either a string or a tokenized text.
        max_char_buffer: Size of buffer that we can run inference on.
        tokenizer_impl: Tokenizer instance to use.
    """

    if isinstance(document, str):
        tokenized_text = tokenizer_impl.tokenize(document)
    else:
        raise ValueError("document has the wrong format. str expected")
    self.tokenized_text = tokenized_text
    self.max_char_buffer = max_char_buffer
    self.sentence_iter = SentenceIterator(self.tokenized_text)
    self.broken_sentence = False
    self.document = document

create_token_interval(start_index, end_index)

Creates a token interval.

Parameters:

Name Type Description Default
start_index int

first token's index (inclusive).

required
end_index int

last token's index + 1 (exclusive).

required

Returns:

Type Description
TokenInterval

Token interval.

Raises:

Type Description
ValueError

If the token indices are invalid.

Source code in src/kibad_llm/extractors/chunking_utils/core.py
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
def create_token_interval(start_index: int, end_index: int) -> tokenizer_lib.TokenInterval:
    """Creates a token interval.

    Args:
      start_index: first token's index (inclusive).
      end_index: last token's index + 1 (exclusive).

    Returns:
      Token interval.

    Raises:
      ValueError: If the token indices are invalid.
    """
    if start_index < 0:
        raise ValueError(f"Start index {start_index} must be positive.")
    if start_index >= end_index:
        raise ValueError(f"Start index {start_index} must be < end index {end_index}.")
    return tokenizer_lib.TokenInterval(start_index=start_index, end_index=end_index)

get_token_interval_text(tokenized_text, token_interval)

Get the text within an interval of tokens.

Parameters:

Name Type Description Default
tokenized_text TokenizedText

Tokenized documents.

required
token_interval TokenInterval

An interval specifying the start (inclusive) and end (exclusive) indices of the tokens to extract. These indices refer to the positions in the list of tokens within tokenized_text.tokens, not the value of the field index of token_pb2.Token. If the tokens are [(index:0, text:A), (index:5, text:B), (index:10, text:C)], we should use token_interval=[0, 2] to represent taking A and B, not [0, 6]. Please see details from the implementation of tokenizer_lib.tokens_text

required

Returns:

Type Description
str

Text within the token interval.

Raises:

Type Description
ValueError

If the token indices are invalid.

TokenUtilError

If tokenizer_lib.tokens_text returns an empty string.

Source code in src/kibad_llm/extractors/chunking_utils/core.py
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
def get_token_interval_text(
    tokenized_text: tokenizer_lib.TokenizedText,
    token_interval: tokenizer_lib.TokenInterval,
) -> str:
    """Get the text within an interval of tokens.

    Args:
      tokenized_text: Tokenized documents.
      token_interval: An interval specifying the start (inclusive) and end
        (exclusive) indices of the tokens to extract. These indices refer to the
        positions in the list of tokens within `tokenized_text.tokens`, not the
        value of the field `index` of `token_pb2.Token`. If the tokens are
        [(index:0, text:A), (index:5, text:B), (index:10, text:C)], we should use
        token_interval=[0, 2] to represent taking A and B, not [0, 6]. Please see
        details from the implementation of tokenizer_lib.tokens_text

    Returns:
      Text within the token interval.

    Raises:
      ValueError: If the token indices are invalid.
      TokenUtilError: If tokenizer_lib.tokens_text returns an empty string.
    """
    if token_interval.start_index >= token_interval.end_index:
        raise ValueError(
            f"Start index {token_interval.start_index} must be < end index "
            f"{token_interval.end_index}."
        )
    return_string = tokenizer_lib.tokens_text(tokenized_text, token_interval)
    logging.debug(
        "Token util returns string: %s for tokenized_text: %s, token_interval:" " %s",
        return_string,
        tokenized_text,
        token_interval,
    )
    if tokenized_text.text and not return_string:
        raise TokenUtilError(
            "Token util returns an empty string unexpectedly. Number of tokens is"
            f" tokenized_text: {len(tokenized_text.tokens)}, token_interval is"
            f" {token_interval.start_index} to {token_interval.end_index}, which"
            " should not lead to empty string."
        )
    return return_string

get_char_interval(tokenized_text, token_interval)

Returns the char interval corresponding to the token interval.

Parameters:

Name Type Description Default
tokenized_text TokenizedText

Document.

required
token_interval TokenInterval

Token interval.

required

Returns:

Type Description
CharInterval

Char interval of the token interval of interest.

Raises:

Type Description
ValueError

If the token_interval is invalid.

Source code in src/kibad_llm/extractors/chunking_utils/core.py
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
def get_char_interval(
    tokenized_text: tokenizer_lib.TokenizedText,
    token_interval: tokenizer_lib.TokenInterval,
) -> CharInterval:
    """Returns the char interval corresponding to the token interval.

    Args:
      tokenized_text: Document.
      token_interval: Token interval.

    Returns:
      Char interval of the token interval of interest.

    Raises:
      ValueError: If the token_interval is invalid.
    """
    if token_interval.start_index >= token_interval.end_index:
        raise ValueError(
            f"Start index {token_interval.start_index} must be < end index "
            f"{token_interval.end_index}."
        )
    start_token = tokenized_text.tokens[token_interval.start_index]
    # Penultimate token prior to interval.end_index
    final_token = tokenized_text.tokens[token_interval.end_index - 1]
    return CharInterval(
        start_pos=start_token.char_interval.start_pos,
        end_pos=final_token.char_interval.end_pos,
    )