Skip to content

F1

F1-based metrics for single fields and collections of fields.

Classes:

Name Description
F1MicroSingleFieldMetric

Compute micro-averaged precision, recall, and F1 for one field.

F1MicroMultipleFieldsMetric

Aggregate single-field F1 metrics across multiple fields.

F1MicroSingleFieldMetric(ignore_missing_entries=False, **kwargs)

Bases: MetricWithTpFpFnEntries

Compute micro-averaged precision, recall, and F1 for one label field.

The metric operates on sets and supports optional field extraction, dictionary flattening, and ignored subfields via the inherited entry-normalization helpers.

Warning

Because the metric compares sets, duplicate predicted labels are collapsed (per record). For example, ["A", "A", "B"] and ["A", "B"] are treated as a perfect match.

See MetricWithPrepareEntryAsSet and MetricWithTpFpFnEntries for keyword arguments for entry-to-set preparation and tp/fp/fn collection.

Source code in src/kibad_llm/metrics/base.py
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
def __init__(self, ignore_missing_entries: bool = False, **kwargs) -> None:
    """Initialize tp/fp/fn entry tracking.

    Args:
        ignore_missing_entries: If `True`, skip updates where either side normalizes to an
            empty set.

    Keyword Args:
        field: Optional field to extract from dictionary inputs.
        flatten_dicts: Whether to flatten nested dictionaries before further processing.
        ignore_subfields: Optional mapping from field names to subfield names that should be
            ignored when converting dictionaries into tuples.
    """
    super().__init__(**kwargs)
    self.ignore_missing_entries = ignore_missing_entries
    self.reset()

calculate_scores(state_counts) staticmethod

Calculate precision, recall, F1, and support from tp/fp/fn counts.

Parameters:

Name Type Description Default
state_counts dict[str, int]

Mapping with the keys tp, fp, and fn.

required

Returns:

Type Description
dict[str, float]

A dictionary containing precision, recall, f1, and support.

Source code in src/kibad_llm/metrics/f1.py
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
@staticmethod
def calculate_scores(state_counts: dict[str, int]) -> dict[str, float]:
    """Calculate precision, recall, F1, and support from tp/fp/fn counts.

    Args:
        state_counts: Mapping with the keys `tp`, `fp`, and `fn`.

    Returns:
        A dictionary containing `precision`, `recall`, `f1`, and `support`.
    """
    tp, fp, fn = state_counts["tp"], state_counts["fp"], state_counts["fn"]
    precision = tp / (tp + fp) if (tp + fp) > 0 else 0.0
    recall = tp / (tp + fn) if (tp + fn) > 0 else 0.0
    f1 = 2 * (precision * recall) / (precision + recall) if (precision + recall) > 0 else 0.0
    return {
        "precision": precision,
        "recall": recall,
        "f1": f1,
        "support": tp + fn,
    }

F1MicroMultipleFieldsMetric(format_as_markdown=True, **kwargs)

Bases: MetricCollectionWithFieldDiscoveryAndGrouping[F1MicroSingleFieldMetric]

Compute single-field F1 scores for multiple fields plus aggregate views.

The metric instantiates one F1MicroSingleFieldMetric per field, optionally expanding nested list/dict fields into generated field names such as organism_trends.Amphibien&Wald. It inherits dynamic field discovery and grouped-field expansion from MetricCollectionWithFieldDiscoveryAndGrouping, and computes additional AVG and ALL aggregate rows on top of the per-field scores.

Attributes:

Name Type Description
fields

Explicit field names to evaluate, or None to discover them dynamically.

format_as_markdown

Whether _format_result should render markdown tables.

subfield_keys

Optional rules for expanding nested dict-like fields into generated fields.

subfield_values

Optional rules restricting which nested values are compared after expansion.

metric_kwargs

Keyword arguments forwarded to the per-field metrics.

Methods:

Name Description
ignore_missing_entries

Expose whether one-sided empty entries are ignored.

Parameters:

Name Type Description Default
format_as_markdown bool

Whether to format the result as a markdown table. Defaults to True.

True

Other Parameters:

Name Type Description
fields

Optional allowlist of fields to evaluate. If omitted, fields are discovered from the union of keys present in each prediction/reference pair.

subfield_keys

Optional mapping describing how nested entries are split into generated fields.

subfield_values

Optional mapping restricting which nested values are kept after field expansion.

sort_fields

Whether to sort the fields in the output. Defaults to False.

flatten_dicts

Whether nested dictionaries should be flattened before comparison.

ignore_subfields

Optional subfields to ignore when hashing dictionary payloads.

ignore_missing_entries

Whether one-sided empty entries should be skipped.

Raises:

Type Description
ValueError

If fields contains the reserved aggregate labels ALL or AVG.

Source code in src/kibad_llm/metrics/f1.py
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
def __init__(
    self,
    format_as_markdown: bool = True,
    **kwargs,
) -> None:
    """Initialize a multi-field F1 metric collection.

    Args:
        format_as_markdown: Whether to format the result as a markdown table. Defaults to True.

    Keyword Args:
        fields: Optional allowlist of fields to evaluate. If omitted, fields are discovered
            from the union of keys present in each prediction/reference pair.
        subfield_keys: Optional mapping describing how nested entries are split into generated
            fields.
        subfield_values: Optional mapping restricting which nested values are kept after field
            expansion.
        sort_fields: Whether to sort the fields in the output. Defaults to False.
        flatten_dicts: Whether nested dictionaries should be flattened before comparison.
        ignore_subfields: Optional subfields to ignore when hashing dictionary payloads.
        ignore_missing_entries: Whether one-sided empty entries should be skipped.

    Raises:
        ValueError: If `fields` contains the reserved aggregate labels `ALL` or `AVG`.
    """
    super().__init__(metric_class=F1MicroSingleFieldMetric, **kwargs)

    # reserve aggregate result field names used by this metric
    if self.fields is not None and ("ALL" in self.fields or "AVG" in self.fields):
        raise ValueError("Fields cannot contain 'ALL' or 'AVG' as field names.")

    self.format_as_markdown = format_as_markdown

ignore_missing_entries property

Return whether one-sided empty entries should be ignored.

Returns:

Type Description
bool

True if per-field metrics skip updates where one side normalizes to an empty set.