Skip to content

Tpfpfn

Collect tp/fp/fn entries, either globally or grouped by record.

Classes:

Name Description
TpFpFnCollector

Return tracked tp/fp/fn entries in JSON-serializable form.

TpFpFnCollectorCollection

Return raw tp/fp/fn entries for multiple fields at once.

TpFpFnCollector(per_record=False, **kwargs)

Bases: MetricWithTpFpFnEntries

Collect tp/fp/fn entries instead of reducing them to scores.

By default, results are returned as JSON-safe [record_id, entry] pairs. With per_record=True, entries are grouped by record via MetricWithTpFpFnEntries.state_per_record.

Attributes:

Name Type Description
per_record

Whether results should be grouped by record instead of returned as global tp/fp/fn lists.

Parameters:

Name Type Description Default
per_record bool

Whether to group results by record id.

False

Other Parameters:

Name Type Description
field

Optional field to extract from dictionary inputs.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary values.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Source code in src/kibad_llm/metrics/tpfpfn.py
26
27
28
29
30
31
32
33
34
35
36
37
38
39
def __init__(self, per_record: bool = False, **kwargs) -> None:
    """Initialize the tp/fp/fn entry collector.

    Args:
        per_record: Whether to group results by record id.

    Keyword Args:
        field: Optional field to extract from dictionary inputs.
        flatten_dicts: Whether nested dictionaries should be flattened before comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary values.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.
    """
    super().__init__(**kwargs)
    self.per_record = per_record

TpFpFnCollectorCollection(**kwargs)

Bases: MetricCollectionWithFieldDiscoveryAndGrouping[TpFpFnCollector]

Collect raw tp/fp/fn entries for multiple fields at once.

The collection lazily creates one TpFpFnCollector per field and inherits optional dynamic field discovery plus grouped-field expansion from MetricCollectionWithFieldDiscoveryAndGrouping. Nested dict-like fields can therefore be expanded into generated field names such as organism_trends.Amphibien&Wald before each per-field collector is updated.

Attributes:

Name Type Description
fields

Explicit field names to evaluate, or None to discover them dynamically.

subfield_keys

Optional rules for expanding nested dict-like fields into generated fields.

subfield_values

Optional rules restricting which nested values are compared after expansion.

metric_kwargs

Keyword arguments forwarded to the per-field TpFpFnCollector instances.

Other Parameters:

Name Type Description
fields

Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair.

subfield_keys

Optional mapping describing how nested entries are split into generated fields.

subfield_values

Optional mapping restricting which nested values are kept after field expansion.

sort_fields

Whether to sort the fields in the output. Defaults to False.

per_record

Whether each per-field collector should group entries by record id.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary payloads.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Source code in src/kibad_llm/metrics/tpfpfn.py
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
def __init__(
    self,
    **kwargs,
) -> None:
    """Initialize a multi-field tp/fp/fn entry collection.

    Keyword Args:
        fields: Optional allowlist of fields to evaluate. If omitted, fields are discovered
            from the union of keys present in each prediction/reference pair.
        subfield_keys: Optional mapping describing how nested entries are split into generated
            fields.
        subfield_values: Optional mapping restricting which nested values are kept after field
            expansion.
        sort_fields: Whether to sort the fields in the output. Defaults to False.
        per_record: Whether each per-field collector should group entries by record id.
        flatten_dicts: Whether nested dictionaries should be flattened before comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary payloads.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.
    """
    super().__init__(metric_class=TpFpFnCollector, **kwargs)