Skip to content

Evaluate

Package entrypoint for dataset evaluation.

Functions:

Name Description
evaluate

Evaluates a dataset containing predictions and references using a specified metric.

main

Helper function to inject the hydra config into evaluate.

evaluate(cfg)

Evaluates a dataset containing predictions and references using a specified metric.

Parameters:

Name Type Description Default
cfg DictConfig

OmegaConf configuration. See configs/evaluate.yaml for details.

required

Returns:

Type Description
dict[str, Any]

A dictionary with evaluation results.

Source code in src/kibad_llm/evaluate.py
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
def evaluate(cfg: DictConfig) -> dict[str, Any]:
    """Evaluates a dataset containing predictions and references using a specified metric.

    Args:
        cfg: OmegaConf configuration. See configs/evaluate.yaml for details.

    Returns:
        A dictionary with evaluation results.
    """
    logger.info("Loading dataset with predictions and references ...")
    logger.info(f"Dataset config: {OmegaConf.to_container(cfg.dataset, resolve=True)}")
    dataset = instantiate(cfg.dataset, _convert_="all")

    logger.info("Instantiating metric ...")
    logger.info(f"Metric config: {OmegaConf.to_container(cfg.metric, resolve=True)}")
    metric: Metric = instantiate(cfg.metric, _convert_="all")

    logger.info("Computing metric ...")
    for record_id, example in dataset.items():
        metric.update(
            prediction=example["prediction"], reference=example["reference"], record_id=record_id
        )
    metric_dict = metric.compute()

    metric.show_result(metric_dict)

    result = {
        RESULT_FORMAT_VERSION_KEY: EVALUATE_VERSION,
        "type": metric.__class__.__name__,
        "data": metric_dict,
    }

    if isinstance(dataset, DictWithMetadata):
        result["prediction"] = dataset.metadata

    return result

main(cfg)

Helper function to inject the hydra config into evaluate.

Parameters:

Name Type Description Default
cfg DictConfig

Evaluation config provided through hydra.

required

Returns:

Type Description
dict[str, Any]

The unchanged output of evaluate

Source code in src/kibad_llm/evaluate.py
80
81
82
83
84
85
86
87
88
89
90
91
92
@hydra.main(
    version_base="1.3", config_path=str(PROJ_ROOT / "configs"), config_name="evaluate.yaml"
)
def main(cfg: DictConfig) -> dict[str, Any]:
    """Helper function to inject the hydra config into [`evaluate`][..evaluate].

    Args:
        cfg: Evaluation config provided through hydra.

    Returns:
        The unchanged output of [`evaluate`][..evaluate]
    """
    return evaluate(cfg)