Skip to content

Metrics

Public metric implementations and metric-related helpers.

Modules:

Name Description
collection

Helpers for grouping multiple metric instances.

confusion_matrix

Confusion-matrix metric based on tp/fp/fn entry tracking for single- and multi-field.

f1

Single-field and multi-field micro-F1 metrics.

errors

Error-collection metrics.

tpfpfn

Raw tp/fp/fn entry collector for single- and multi-field.

Classes:

Name Description
MetricCollection

Aggregate multiple sub-metrics.

ConfusionMatrix

Build confusion matrices from tp/fp/fn entry state for one field.

ConfusionMatrixCollection

Build confusion matrices from tp/fp/fn entry state for multiple fields at once.

F1MicroSingleFieldMetric

Compute micro-F1 for one field.

F1MicroMultipleFieldsMetric

Compute micro-F1 for multiple fields plus aggregates.

ErrorCollector

Collect and count prediction errors.

TpFpFnCollector

Return raw tp/fp/fn entries for inspection.

TpFpFnCollectorCollection

Return raw tp/fp/fn entries for multiple fields at once.

MetricCollection(metrics=None, sort_fields=False)

Bases: Metric, Generic[T]

A metric that aggregates multiple sub-metrics.

Parameters:

Name Type Description Default
metrics dict[str, T] | None

Optional mapping of metric names to metric instances.

None
sort_fields bool

Whether computed results should be emitted in sorted field order.

False
Source code in src/kibad_llm/metrics/collection.py
25
26
27
28
29
30
31
32
33
34
def __init__(self, metrics: dict[str, T] | None = None, sort_fields: bool = False) -> None:
    """Initialize the metric collection.

    Args:
        metrics: Optional mapping of metric names to metric instances.
        sort_fields: Whether computed results should be emitted in sorted field order.
    """
    super().__init__()
    self.metrics: dict[str, T] = metrics or dict()
    self.sort_fields = sort_fields

add_metric(name, metric)

Adds a new metric to the collection.

Source code in src/kibad_llm/metrics/collection.py
36
37
38
39
40
def add_metric(self, name: str, metric: T) -> None:
    """Adds a new metric to the collection."""
    if name in self.metrics:
        raise ValueError(f"Metric {name} already exists")
    self.metrics[name] = metric

reset()

Resets all sub-metrics.

Source code in src/kibad_llm/metrics/collection.py
42
43
44
45
def reset(self) -> None:
    """Resets all sub-metrics."""
    for metric in self.metrics.values():
        metric.reset()

ConfusionMatrix(show_as_markdown=False, unassignable_label='UNASSIGNABLE', undetected_label='UNDETECTED', **kwargs)

Bases: MetricWithTpFpFnEntries

Build an alignment-averaged confusion matrix from tp/fp/fn entry state.

In multi-label settings, unmatched gold and predicted labels do not define a unique off-diagonal confusion matrix. For example, if one record contains a missed gold label A and an extra predicted label B, the tp/fp/fn state alone does not tell us whether this should be accounted for as A -> B, as A -> UNDETECTED plus UNASSIGNABLE -> B, or as part of another possible alignment.

This metric therefore uses an alignment-averaged accounting rule per record:

  • exact true-positive labels are counted deterministically on the diagonal;
  • unmatched gold and predicted labels are distributed over all compatible partial one-to-one alignments between false negatives and false positives;
  • all compatible partial alignments are weighted equally;
  • unmatched gold labels that are not aligned to a prediction contribute to undetected_label;
  • unmatched predicted labels that are not aligned to a gold label contribute to unassignable_label.

The resulting off-diagonal entries are expected ambiguous error mass under this uniform partial-alignment assumption. They are useful for exploratory error analysis, but they are not directly observed misclassification counts and do not provide statistical significance or uncertainty estimates.

Warning

Because the metric operates on sets, duplicate predicted or gold labels are collapsed in multi-label settings per record.

Warning

Off-diagonal entries indicate possible label-shift mass under the accounting rule, not evidence that a particular label shift actually occurred.

Parameters:

Name Type Description Default
show_as_markdown bool

Whether compute() should log the resulting confusion matrix as a markdown table.

False
unassignable_label str

Label used on the gold axis for predicted labels that remain unaligned to any gold label.

'UNASSIGNABLE'
undetected_label str

Label used on the prediction axis for gold labels that remain unaligned to any prediction.

'UNDETECTED'

Other Parameters:

Name Type Description
field

Optional field to extract from dictionary inputs.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary values.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Source code in src/kibad_llm/metrics/confusion_matrix.py
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
def __init__(
    self,
    show_as_markdown: bool = False,
    unassignable_label: str = "UNASSIGNABLE",
    undetected_label: str = "UNDETECTED",
    **kwargs: Any,
):
    """Initialize the confusion-matrix metric.

    Args:
        show_as_markdown: Whether `compute()` should log the resulting confusion
            matrix as a markdown table.
        unassignable_label: Label used on the gold axis for predicted labels that
            remain unaligned to any gold label.
        undetected_label: Label used on the prediction axis for gold labels that
            remain unaligned to any prediction.

    Keyword Args:
        field: Optional field to extract from dictionary inputs.
        flatten_dicts: Whether nested dictionaries should be flattened before
            comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary
            values.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.
    """
    super().__init__(**kwargs)
    self.unassignable_label = unassignable_label
    self.undetected_label = undetected_label
    self.show_as_markdown = show_as_markdown

ConfusionMatrixCollection(**kwargs)

Bases: MetricCollectionWithFieldDiscoveryAndGrouping[ConfusionMatrix]

Build confusion matrices for multiple fields at once.

The collection lazily creates one ConfusionMatrix per field and inherits optional dynamic field discovery plus grouped-field expansion from MetricCollectionWithFieldDiscoveryAndGrouping. Nested dict-like fields can therefore be expanded into generated field names such as organism_trends.Amphibien&Wald before each per-field confusion matrix is updated.

Attributes:

Name Type Description
fields

Explicit field names to evaluate, or None to discover them dynamically.

subfield_keys

Optional rules for expanding nested dict-like fields into generated fields.

subfield_values

Optional rules restricting which nested values are compared after expansion.

metric_kwargs

Keyword arguments forwarded to the per-field ConfusionMatrix instances.

Other Parameters:

Name Type Description
fields

Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair.

subfield_keys

Optional mapping describing how nested entries are split into generated fields.

subfield_values

Optional mapping restricting which nested values are kept after field expansion.

sort_fields

Whether to sort the fields in the output. Defaults to False.

show_as_markdown

Whether each per-field confusion matrix should be logged as a markdown table when computed.

unassignable_label

Label used on the gold axis for false positives.

undetected_label

Label used on the prediction axis for false negatives.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary payloads.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Source code in src/kibad_llm/metrics/confusion_matrix.py
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
def __init__(
    self,
    **kwargs,
) -> None:
    """Initialize a multi-field confusion-matrix collection.

    Keyword Args:
        fields: Optional allowlist of fields to evaluate. If omitted, fields are discovered
            from the union of keys present in each prediction/reference pair.
        subfield_keys: Optional mapping describing how nested entries are split into generated
            fields.
        subfield_values: Optional mapping restricting which nested values are kept after field
            expansion.
        sort_fields: Whether to sort the fields in the output. Defaults to False.
        show_as_markdown: Whether each per-field confusion matrix should be logged as a markdown
            table when computed.
        unassignable_label: Label used on the gold axis for false positives.
        undetected_label: Label used on the prediction axis for false negatives.
        flatten_dicts: Whether nested dictionaries should be flattened before comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary payloads.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.
    """
    super().__init__(metric_class=ConfusionMatrix, **kwargs)

ErrorCollector(show_errors=False, type_separator=': ')

Bases: Metric

Collect error messages from model predictions and summarize their occurrences.

Parameters:

Name Type Description Default
show_errors bool

If True, log each collected error message.

False
type_separator str

Separator used to split error messages into type and details.

': '

Parameters:

Name Type Description Default
show_errors bool

If True, log each collected error message as it is collected.

False
type_separator str

Separator used to split error messages into type and details.

': '
Source code in src/kibad_llm/metrics/errors.py
25
26
27
28
29
30
31
32
33
34
def __init__(self, show_errors: bool = False, type_separator: str = ": ") -> None:
    """Initialize the error collector.

    Args:
        show_errors: If `True`, log each collected error message as it is collected.
        type_separator: Separator used to split error messages into type and details.
    """
    self.show_errors = show_errors
    self.type_separator = type_separator
    self.reset()

reset()

Reset the collected error state.

Source code in src/kibad_llm/metrics/errors.py
36
37
38
def reset(self) -> None:
    """Reset the collected error state."""
    self.state: list[list[str]] = []

F1MicroMultipleFieldsMetric(format_as_markdown=True, **kwargs)

Bases: MetricCollectionWithFieldDiscoveryAndGrouping[F1MicroSingleFieldMetric]

Compute single-field F1 scores for multiple fields plus aggregate views.

The metric instantiates one F1MicroSingleFieldMetric per field, optionally expanding nested list/dict fields into generated field names such as organism_trends.Amphibien&Wald. It inherits dynamic field discovery and grouped-field expansion from MetricCollectionWithFieldDiscoveryAndGrouping, and computes additional AVG and ALL aggregate rows on top of the per-field scores.

Attributes:

Name Type Description
fields

Explicit field names to evaluate, or None to discover them dynamically.

format_as_markdown

Whether _format_result should render markdown tables.

subfield_keys

Optional rules for expanding nested dict-like fields into generated fields.

subfield_values

Optional rules restricting which nested values are compared after expansion.

metric_kwargs

Keyword arguments forwarded to the per-field metrics.

Methods:

Name Description
ignore_missing_entries

Expose whether one-sided empty entries are ignored.

Parameters:

Name Type Description Default
format_as_markdown bool

Whether to format the result as a markdown table. Defaults to True.

True

Other Parameters:

Name Type Description
fields

Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair.

subfield_keys

Optional mapping describing how nested entries are split into generated fields.

subfield_values

Optional mapping restricting which nested values are kept after field expansion.

sort_fields

Whether to sort the fields in the output. Defaults to False.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary payloads.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Raises:

Type Description
ValueError

If fields contains the reserved aggregate labels ALL or AVG.

Source code in src/kibad_llm/metrics/f1.py
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
def __init__(
    self,
    format_as_markdown: bool = True,
    **kwargs,
) -> None:
    """Initialize a multi-field F1 metric collection.

    Args:
        format_as_markdown: Whether to format the result as a markdown table. Defaults to True.

    Keyword Args:
        fields: Optional allowlist of fields to evaluate. If omitted, fields are discovered
            from the union of keys present in each prediction/reference pair.
        subfield_keys: Optional mapping describing how nested entries are split into generated
            fields.
        subfield_values: Optional mapping restricting which nested values are kept after field
            expansion.
        sort_fields: Whether to sort the fields in the output. Defaults to False.
        flatten_dicts: Whether nested dictionaries should be flattened before comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary payloads.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.

    Raises:
        ValueError: If `fields` contains the reserved aggregate labels `ALL` or `AVG`.
    """
    super().__init__(metric_class=F1MicroSingleFieldMetric, **kwargs)

    # reserve aggregate result field names used by this metric
    if self.fields is not None and ("ALL" in self.fields or "AVG" in self.fields):
        raise ValueError("Fields cannot contain 'ALL' or 'AVG' as field names.")

    self.format_as_markdown = format_as_markdown

ignore_missing_entries property

Return whether one-sided empty entries should be ignored.

Returns:

Type Description
bool

True if per-field metrics skip updates where one side normalizes to an empty set.

F1MicroSingleFieldMetric(ignore_missing_entries=False, **kwargs)

Bases: MetricWithTpFpFnEntries

Compute micro-averaged precision, recall, and F1 for one label field.

The metric operates on sets and supports optional field extraction, dictionary flattening, and ignored subfields via the inherited entry-normalization helpers.

Warning

Because the metric compares sets, duplicate predicted labels are collapsed (per record). For example, ["A", "A", "B"] and ["A", "B"] are treated as a perfect match.

See MetricWithPrepareEntryAsSet and MetricWithTpFpFnEntries for keyword arguments for entry-to-set preparation and tp/fp/fn collection.

Source code in src/kibad_llm/metrics/base.py
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
def __init__(self, ignore_missing_entries: bool = False, **kwargs) -> None:
    """Initialize tp/fp/fn entry tracking.

    Args:
        ignore_missing_entries: If `True`, skip updates where either side normalizes to an
            empty set.

    Keyword Args:
        field: Optional field to extract from dictionary inputs.
        flatten_dicts: Whether to flatten nested dictionaries before further processing.
        ignore_subfields: Optional mapping from field names to subfield names that should be
            ignored when converting dictionaries into tuples.
    """
    super().__init__(**kwargs)
    self.ignore_missing_entries = ignore_missing_entries
    self.reset()

calculate_scores(state_counts) staticmethod

Calculate precision, recall, F1, and support from tp/fp/fn counts.

Parameters:

Name Type Description Default
state_counts dict[str, int]

Mapping with the keys tp, fp, and fn.

required

Returns:

Type Description
dict[str, float]

A dictionary containing precision, recall, f1, and support.

Source code in src/kibad_llm/metrics/f1.py
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
@staticmethod
def calculate_scores(state_counts: dict[str, int]) -> dict[str, float]:
    """Calculate precision, recall, F1, and support from tp/fp/fn counts.

    Args:
        state_counts: Mapping with the keys `tp`, `fp`, and `fn`.

    Returns:
        A dictionary containing `precision`, `recall`, `f1`, and `support`.
    """
    tp, fp, fn = state_counts["tp"], state_counts["fp"], state_counts["fn"]
    precision = tp / (tp + fp) if (tp + fp) > 0 else 0.0
    recall = tp / (tp + fn) if (tp + fn) > 0 else 0.0
    f1 = 2 * (precision * recall) / (precision + recall) if (precision + recall) > 0 else 0.0
    return {
        "precision": precision,
        "recall": recall,
        "f1": f1,
        "support": tp + fn,
    }

TpFpFnCollector(per_record=False, **kwargs)

Bases: MetricWithTpFpFnEntries

Collect tp/fp/fn entries instead of reducing them to scores.

By default, results are returned as JSON-safe [record_id, entry] pairs. With per_record=True, entries are grouped by record via MetricWithTpFpFnEntries.state_per_record.

Attributes:

Name Type Description
per_record

Whether results should be grouped by record instead of returned as global tp/fp/fn lists.

Parameters:

Name Type Description Default
per_record bool

Whether to group results by record id.

False

Other Parameters:

Name Type Description
field

Optional field to extract from dictionary inputs.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary values.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Source code in src/kibad_llm/metrics/tpfpfn.py
26
27
28
29
30
31
32
33
34
35
36
37
38
39
def __init__(self, per_record: bool = False, **kwargs) -> None:
    """Initialize the tp/fp/fn entry collector.

    Args:
        per_record: Whether to group results by record id.

    Keyword Args:
        field: Optional field to extract from dictionary inputs.
        flatten_dicts: Whether nested dictionaries should be flattened before comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary values.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.
    """
    super().__init__(**kwargs)
    self.per_record = per_record

TpFpFnCollectorCollection(**kwargs)

Bases: MetricCollectionWithFieldDiscoveryAndGrouping[TpFpFnCollector]

Collect raw tp/fp/fn entries for multiple fields at once.

The collection lazily creates one TpFpFnCollector per field and inherits optional dynamic field discovery plus grouped-field expansion from MetricCollectionWithFieldDiscoveryAndGrouping. Nested dict-like fields can therefore be expanded into generated field names such as organism_trends.Amphibien&Wald before each per-field collector is updated.

Attributes:

Name Type Description
fields

Explicit field names to evaluate, or None to discover them dynamically.

subfield_keys

Optional rules for expanding nested dict-like fields into generated fields.

subfield_values

Optional rules restricting which nested values are compared after expansion.

metric_kwargs

Keyword arguments forwarded to the per-field TpFpFnCollector instances.

Other Parameters:

Name Type Description
fields

Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair.

subfield_keys

Optional mapping describing how nested entries are split into generated fields.

subfield_values

Optional mapping restricting which nested values are kept after field expansion.

sort_fields

Whether to sort the fields in the output. Defaults to False.

per_record

Whether each per-field collector should group entries by record id.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary payloads.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Source code in src/kibad_llm/metrics/tpfpfn.py
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
def __init__(
    self,
    **kwargs,
) -> None:
    """Initialize a multi-field tp/fp/fn entry collection.

    Keyword Args:
        fields: Optional allowlist of fields to evaluate. If omitted, fields are discovered
            from the union of keys present in each prediction/reference pair.
        subfield_keys: Optional mapping describing how nested entries are split into generated
            fields.
        subfield_values: Optional mapping restricting which nested values are kept after field
            expansion.
        sort_fields: Whether to sort the fields in the output. Defaults to False.
        per_record: Whether each per-field collector should group entries by record id.
        flatten_dicts: Whether nested dictionaries should be flattened before comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary payloads.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.
    """
    super().__init__(metric_class=TpFpFnCollector, **kwargs)