Metrics
Public metric implementations and metric-related helpers.
Modules:
| Name | Description |
|---|---|
collection |
Helpers for grouping multiple metric instances. |
confusion_matrix |
Confusion-matrix metric based on tp/fp/fn entry tracking for single- and multi-field. |
f1 |
Single-field and multi-field micro-F1 metrics. |
errors |
Error-collection metrics. |
tpfpfn |
Raw tp/fp/fn entry collector for single- and multi-field. |
Classes:
| Name | Description |
|---|---|
MetricCollection |
Aggregate multiple sub-metrics. |
ConfusionMatrix |
Build confusion matrices from tp/fp/fn entry state for one field. |
ConfusionMatrixCollection |
Build confusion matrices from tp/fp/fn entry state for multiple fields at once. |
F1MicroSingleFieldMetric |
Compute micro-F1 for one field. |
F1MicroMultipleFieldsMetric |
Compute micro-F1 for multiple fields plus aggregates. |
ErrorCollector |
Collect and count prediction errors. |
TpFpFnCollector |
Return raw tp/fp/fn entries for inspection. |
TpFpFnCollectorCollection |
Return raw tp/fp/fn entries for multiple fields at once. |
MetricCollection(metrics=None, sort_fields=False)
Bases: Metric, Generic[T]
A metric that aggregates multiple sub-metrics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
metrics
|
dict[str, T] | None
|
Optional mapping of metric names to metric instances. |
None
|
sort_fields
|
bool
|
Whether computed results should be emitted in sorted field order. |
False
|
Source code in src/kibad_llm/metrics/collection.py
25 26 27 28 29 30 31 32 33 34 | |
add_metric(name, metric)
Adds a new metric to the collection.
Source code in src/kibad_llm/metrics/collection.py
36 37 38 39 40 | |
reset()
Resets all sub-metrics.
Source code in src/kibad_llm/metrics/collection.py
42 43 44 45 | |
ConfusionMatrix(show_as_markdown=False, unassignable_label='UNASSIGNABLE', undetected_label='UNDETECTED', **kwargs)
Bases: MetricWithTpFpFnEntries
Build an alignment-averaged confusion matrix from tp/fp/fn entry state.
In multi-label settings, unmatched gold and predicted labels do not define a unique
off-diagonal confusion matrix. For example, if one record contains a missed gold
label A and an extra predicted label B, the tp/fp/fn state alone does not tell
us whether this should be accounted for as A -> B, as A -> UNDETECTED plus
UNASSIGNABLE -> B, or as part of another possible alignment.
This metric therefore uses an alignment-averaged accounting rule per record:
- exact true-positive labels are counted deterministically on the diagonal;
- unmatched gold and predicted labels are distributed over all compatible partial one-to-one alignments between false negatives and false positives;
- all compatible partial alignments are weighted equally;
- unmatched gold labels that are not aligned to a prediction contribute to
undetected_label; - unmatched predicted labels that are not aligned to a gold label contribute to
unassignable_label.
The resulting off-diagonal entries are expected ambiguous error mass under this uniform partial-alignment assumption. They are useful for exploratory error analysis, but they are not directly observed misclassification counts and do not provide statistical significance or uncertainty estimates.
Warning
Because the metric operates on sets, duplicate predicted or gold labels are collapsed in multi-label settings per record.
Warning
Off-diagonal entries indicate possible label-shift mass under the accounting rule, not evidence that a particular label shift actually occurred.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
show_as_markdown
|
bool
|
Whether |
False
|
unassignable_label
|
str
|
Label used on the gold axis for predicted labels that remain unaligned to any gold label. |
'UNASSIGNABLE'
|
undetected_label
|
str
|
Label used on the prediction axis for gold labels that remain unaligned to any prediction. |
'UNDETECTED'
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
field |
Optional field to extract from dictionary inputs. |
|
flatten_dicts |
Whether nested dictionaries should be flattened before comparison. |
|
ignore_subfields |
Optional subfields to ignore when hashing dictionary values. |
|
ignore_missing_entries |
Whether one-sided empty entries should be skipped. |
Source code in src/kibad_llm/metrics/confusion_matrix.py
76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 | |
ConfusionMatrixCollection(**kwargs)
Bases: MetricCollectionWithFieldDiscoveryAndGrouping[ConfusionMatrix]
Build confusion matrices for multiple fields at once.
The collection lazily creates one ConfusionMatrix per field and inherits optional dynamic
field discovery plus grouped-field expansion from
MetricCollectionWithFieldDiscoveryAndGrouping. Nested dict-like fields can therefore be
expanded into generated field names such as organism_trends.Amphibien&Wald before each
per-field confusion matrix is updated.
Attributes:
| Name | Type | Description |
|---|---|---|
fields |
Explicit field names to evaluate, or |
|
subfield_keys |
Optional rules for expanding nested dict-like fields into generated fields. |
|
subfield_values |
Optional rules restricting which nested values are compared after expansion. |
|
metric_kwargs |
Keyword arguments forwarded to the per-field |
Other Parameters:
| Name | Type | Description |
|---|---|---|
fields |
Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair. |
|
subfield_keys |
Optional mapping describing how nested entries are split into generated fields. |
|
subfield_values |
Optional mapping restricting which nested values are kept after field expansion. |
|
sort_fields |
Whether to sort the fields in the output. Defaults to False. |
|
show_as_markdown |
Whether each per-field confusion matrix should be logged as a markdown table when computed. |
|
unassignable_label |
Label used on the gold axis for false positives. |
|
undetected_label |
Label used on the prediction axis for false negatives. |
|
flatten_dicts |
Whether nested dictionaries should be flattened before comparison. |
|
ignore_subfields |
Optional subfields to ignore when hashing dictionary payloads. |
|
ignore_missing_entries |
Whether one-sided empty entries should be skipped. |
Source code in src/kibad_llm/metrics/confusion_matrix.py
274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 | |
ErrorCollector(show_errors=False, type_separator=': ')
Bases: Metric
Collect error messages from model predictions and summarize their occurrences.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
show_errors
|
bool
|
If |
False
|
type_separator
|
str
|
Separator used to split error messages into type and details. |
': '
|
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
show_errors
|
bool
|
If |
False
|
type_separator
|
str
|
Separator used to split error messages into type and details. |
': '
|
Source code in src/kibad_llm/metrics/errors.py
25 26 27 28 29 30 31 32 33 34 | |
reset()
Reset the collected error state.
Source code in src/kibad_llm/metrics/errors.py
36 37 38 | |
F1MicroMultipleFieldsMetric(format_as_markdown=True, **kwargs)
Bases: MetricCollectionWithFieldDiscoveryAndGrouping[F1MicroSingleFieldMetric]
Compute single-field F1 scores for multiple fields plus aggregate views.
The metric instantiates one F1MicroSingleFieldMetric per field, optionally expanding nested
list/dict fields into generated field names such as organism_trends.Amphibien&Wald. It
inherits dynamic field discovery and grouped-field expansion from
MetricCollectionWithFieldDiscoveryAndGrouping, and computes additional AVG and ALL
aggregate rows on top of the per-field scores.
Attributes:
| Name | Type | Description |
|---|---|---|
fields |
Explicit field names to evaluate, or |
|
format_as_markdown |
Whether |
|
subfield_keys |
Optional rules for expanding nested dict-like fields into generated fields. |
|
subfield_values |
Optional rules restricting which nested values are compared after expansion. |
|
metric_kwargs |
Keyword arguments forwarded to the per-field metrics. |
Methods:
| Name | Description |
|---|---|
ignore_missing_entries |
Expose whether one-sided empty entries are ignored. |
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
format_as_markdown
|
bool
|
Whether to format the result as a markdown table. Defaults to True. |
True
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
fields |
Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair. |
|
subfield_keys |
Optional mapping describing how nested entries are split into generated fields. |
|
subfield_values |
Optional mapping restricting which nested values are kept after field expansion. |
|
sort_fields |
Whether to sort the fields in the output. Defaults to False. |
|
flatten_dicts |
Whether nested dictionaries should be flattened before comparison. |
|
ignore_subfields |
Optional subfields to ignore when hashing dictionary payloads. |
|
ignore_missing_entries |
Whether one-sided empty entries should be skipped. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/kibad_llm/metrics/f1.py
84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 | |
ignore_missing_entries
property
Return whether one-sided empty entries should be ignored.
Returns:
| Type | Description |
|---|---|
bool
|
|
F1MicroSingleFieldMetric(ignore_missing_entries=False, **kwargs)
Bases: MetricWithTpFpFnEntries
Compute micro-averaged precision, recall, and F1 for one label field.
The metric operates on sets and supports optional field extraction, dictionary flattening, and ignored subfields via the inherited entry-normalization helpers.
Warning
Because the metric compares sets, duplicate predicted labels are collapsed (per record).
For example, ["A", "A", "B"] and ["A", "B"] are treated as a perfect match.
See MetricWithPrepareEntryAsSet and MetricWithTpFpFnEntries for keyword arguments
for entry-to-set preparation and tp/fp/fn collection.
Source code in src/kibad_llm/metrics/base.py
185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 | |
calculate_scores(state_counts)
staticmethod
Calculate precision, recall, F1, and support from tp/fp/fn counts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
state_counts
|
dict[str, int]
|
Mapping with the keys |
required |
Returns:
| Type | Description |
|---|---|
dict[str, float]
|
A dictionary containing |
Source code in src/kibad_llm/metrics/f1.py
31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 | |
TpFpFnCollector(per_record=False, **kwargs)
Bases: MetricWithTpFpFnEntries
Collect tp/fp/fn entries instead of reducing them to scores.
By default, results are returned as JSON-safe [record_id, entry] pairs. With
per_record=True, entries are grouped by record via
MetricWithTpFpFnEntries.state_per_record.
Attributes:
| Name | Type | Description |
|---|---|---|
per_record |
Whether results should be grouped by record instead of returned as global tp/fp/fn lists. |
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
per_record
|
bool
|
Whether to group results by record id. |
False
|
Other Parameters:
| Name | Type | Description |
|---|---|---|
field |
Optional field to extract from dictionary inputs. |
|
flatten_dicts |
Whether nested dictionaries should be flattened before comparison. |
|
ignore_subfields |
Optional subfields to ignore when hashing dictionary values. |
|
ignore_missing_entries |
Whether one-sided empty entries should be skipped. |
Source code in src/kibad_llm/metrics/tpfpfn.py
26 27 28 29 30 31 32 33 34 35 36 37 38 39 | |
TpFpFnCollectorCollection(**kwargs)
Bases: MetricCollectionWithFieldDiscoveryAndGrouping[TpFpFnCollector]
Collect raw tp/fp/fn entries for multiple fields at once.
The collection lazily creates one TpFpFnCollector per field and inherits optional dynamic
field discovery plus grouped-field expansion from
MetricCollectionWithFieldDiscoveryAndGrouping. Nested dict-like fields can therefore be
expanded into generated field names such as organism_trends.Amphibien&Wald before each
per-field collector is updated.
Attributes:
| Name | Type | Description |
|---|---|---|
fields |
Explicit field names to evaluate, or |
|
subfield_keys |
Optional rules for expanding nested dict-like fields into generated fields. |
|
subfield_values |
Optional rules restricting which nested values are compared after expansion. |
|
metric_kwargs |
Keyword arguments forwarded to the per-field |
Other Parameters:
| Name | Type | Description |
|---|---|---|
fields |
Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair. |
|
subfield_keys |
Optional mapping describing how nested entries are split into generated fields. |
|
subfield_values |
Optional mapping restricting which nested values are kept after field expansion. |
|
sort_fields |
Whether to sort the fields in the output. Defaults to False. |
|
per_record |
Whether each per-field collector should group entries by record id. |
|
flatten_dicts |
Whether nested dictionaries should be flattened before comparison. |
|
ignore_subfields |
Optional subfields to ignore when hashing dictionary payloads. |
|
ignore_missing_entries |
Whether one-sided empty entries should be skipped. |
Source code in src/kibad_llm/metrics/tpfpfn.py
90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 | |