Skip to content

Hydra callbacks

Saving of Hydra data.

Modules:

Name Description
save_job_return_value

Provides a class for the saving of Hydra jobs' and multiruns' outputs.

Classes:

Name Description
SaveJobReturnValueCallback

Handles the saving of job return-values at the ends of jobs and multiruns.

SaveJobReturnValueCallback(filenames='job_return_value.json', integrate_multirun_result=False, multirun_aggregator_blacklist=None, sort_markdown_columns=True, markdown_round_digits=3, markdown_data_key=None, multirun_create_ids_from_overrides=True, multirun_job_id_key='job_id', multirun_convert_job_ids=False, handle_previous_result=None, replace_existing_overrides=False, multirun_add_overrides_as_dict=False, multirun_show_file_contents=None, multirun_overrides_separator='-', multirun_markdown_group_by=None, multirun_markdown_transpose=False, paths_file=None, path_id=None, multirun_paths_file=None, multirun_path_id=None)

Bases: Callback

Save each job's return-value in ${output_dir}/${filename}, for every entry in filenames.

This also works for multi-runs (e.g. sweeps for hyperparameter search). In this case, the result will be saved additionally in a common file in the multi-run log directory. If integrate_multirun_result=True, the job return-values are also aggregated (e.g. mean, min, max) and saved in another file.

This class exists to postprocess job and multirun outputs, by overwriting Hydra's no-op callback hooks.

Outputs can be saved as json and markdown. The markdown output is just a table with no text around it.

For more info on the Attributes, refer to their __init__ args counterparts.

Attributes:

Name Type Description
_log Logger

The logger of this callback.

_job_returns list[JobReturn]

The return objects of all jobs seen by on_job_end so far. Consumed by on_multirun_end.

filenames list[str]

Used by on_job_end:

Attributes:

Name Type Description
handle_previous_result str | None
replace_existing_overrides bool
paths_file str | None
markdown_data_key str | None

Used by on_multirun_end:

Attributes:

Name Type Description
integrate_multirun_result bool
multirun_aggregator_blacklist list[str] | None
multirun_create_ids_from_overrides bool
multirun_job_id_key str
multirun_convert_job_ids bool
multirun_add_overrides_as_dict bool
multirun_show_file_contents list[str]
multirun_overrides_separator str
multirun_markdown_group_by list[str] | None
multirun_paths_file str | None
multirun_path_id str | None

Used by _save:

Attributes:

Name Type Description
sort_markdown_columns bool
markdown_round_digits int | None
multirun_markdown_transpose bool

Not used:

Attributes:

Name Type Description
path_id str | None

Methods:

Name Description
on_job_end

Save a single job's return-value once the job finishes.

on_multirun_end

Collate and save all jobs' return-values once the multi-run finishes.

Parameters:

Name Type Description Default
filenames str | list[str]

The filename(s) of the file(s) to save the job return-value to. If a string is passed, it will be wrapped in a list. The internal type hence is always a list of strings. If an entry ends with ".json", the return-value is saved as a json file. If it ends with ".md", the return-value is saved as a markdown file. Json files are more complete data wise, whilst markdown files have more settings that can be applied for readability.

'job_return_value.json'
integrate_multirun_result bool

If True, the job return-values of all jobs from a multi-run are rearranged into a dict of lists (maybe nested), where the keys are the keys of the job return-values and the values are lists of the corresponding values of all jobs. This is useful if you want to access specific values of all jobs in a multi-run all at once. Also, aggregated values (e.g. mean, min, max) are created for all numeric values and saved in another file.

False
multirun_aggregator_blacklist list[str] | None

A list of keys to exclude from the aggregation (of multirun results), such as "count" or "25%". If None, all keys are included. See pd.DataFrame.describe() for possible aggregation keys. For numeric values, it is recommended to use ["min", "25%", "50%", "75%", "max"] which will result in keeping only the count, mean and std values.

None
sort_markdown_columns bool

If True, the columns of the markdown table are sorted alphabetically.

True
markdown_round_digits int | None

The number of digits to round the values in the markdown file to. If None, no rounding is applied.

3
markdown_data_key str | None

If provided and present in the job return-value, save only the value at this key when saving single job results to markdown. This is useful to strip metadata from the job result and, thus, allow for correct table formatting of metric results, for instance. If the key is absent (or None), a legacy path is used instead, which drops the handle_previous_result field and the result format version key from the markdown output.

None
multirun_create_ids_from_overrides bool

If True, create job identifiers from the overrides of the jobs in a multi-run. If False, the job index is used as identifier.

True
multirun_job_id_key str

The key to use for the job identifiers in the integrated multi-run result.

'job_id'
multirun_convert_job_ids bool

If True, convert job ids to dictionaries. Works only if integrate_multirun_result is True.

False
handle_previous_result str | None

If provided, assume the job return-value contains a field with the given name (e.g. "prediction") that itself contains an "overrides" field. The overrides from this field are either used to replace the existing overrides in the job return object (if replace_existing_overrides is True) or simply converted to a dictionary and added back to the job return-value (if replace_existing_overrides is False). Furthermore, on the legacy markdown path (see markdown_data_key), the field is removed from the job return-value before saving it as markdown to avoid destroying the table structure.

None
replace_existing_overrides bool

If True, replace existing overrides in the job return-value with the overrides from the job return object if available. If False, the overrides are just converted to a dictionary, if available.

False
multirun_add_overrides_as_dict bool

If True, add the overrides as a dictionary to each job return-value under the key "overrides".

False
multirun_show_file_contents list[str] | None

A list of filenames (from the filenames attribute or aggregated files) whose contents are logged to the console after saving the multi-run results. If None is passed, saves [] instead.

None
multirun_overrides_separator str

The separator to use when creating job identifiers from overrides.

'-'
multirun_markdown_group_by str | list[str] | None

The column(s) to group by when saving the multi-run result as a markdown file. For numeric columns, the mean and std are calculated. For non-numeric columns, a list of values is created. If None, no grouping is applied. A single string is wrapped into a list. If a string is passed, it will be wrapped in a list.

None
multirun_markdown_transpose bool

If True, transpose the markdown table for multi-run results.

False
paths_file str | None

The file to append the paths of the log directories to. If None, the paths are not saved.

None
path_id str | None

This is currently not in use. A prefix to add to each line in the paths_file, separated by a colon. If None, no prefix is added.

None
multirun_paths_file str | None

The file to save the paths of the multi-run log directories to. If None, the paths are not saved.

None
multirun_path_id str | None

A prefix to add to each line in the multirun_paths_file, separated by a colon. If None, no prefix is added.

None
Source code in src/kibad_llm/hydra_callbacks/save_job_return_value.py
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
def __init__(
    self,
    filenames: str | list[str] = "job_return_value.json",
    integrate_multirun_result: bool = False,
    multirun_aggregator_blacklist: list[str] | None = None,
    sort_markdown_columns: bool = True,
    markdown_round_digits: int | None = 3,
    markdown_data_key: str | None = None,
    multirun_create_ids_from_overrides: bool = True,
    multirun_job_id_key: str = "job_id",
    multirun_convert_job_ids: bool = False,
    handle_previous_result: str | None = None,
    replace_existing_overrides: bool = False,
    multirun_add_overrides_as_dict: bool = False,
    multirun_show_file_contents: list[str] | None = None,
    multirun_overrides_separator: str = "-",
    multirun_markdown_group_by: str | list[str] | None = None,
    multirun_markdown_transpose: bool = False,
    paths_file: str | None = None,
    path_id: str | None = None,
    multirun_paths_file: str | None = None,
    multirun_path_id: str | None = None,
) -> None:
    """Assign args to attributes and do some safety conversions beforehand.

    Args:
        filenames: The filename(s) of the file(s) to save the job return-value to. If a string is passed, it will
            be wrapped in a list. The internal type hence is always a list of strings.
            If an entry ends with ".json", the return-value is saved as a json file. If it ends
            with ".md", the return-value is saved as a markdown file. Json files are more complete data wise, whilst
            markdown files have more settings that can be applied for readability.
        integrate_multirun_result: If True, the job return-values of all jobs from a multi-run are rearranged
            into a dict of lists (maybe nested), where the keys are the keys of the job return-values and the values
            are lists of the corresponding values of all jobs. This is useful if you want to access specific values
            of all jobs in a multi-run all at once. Also, aggregated values (e.g. mean, min, max) are created for all
            numeric values and saved in another file.
        multirun_aggregator_blacklist: A list of keys to exclude from the aggregation (of multirun
            results), such as "count" or "25%". If None, all keys are included. See `pd.DataFrame.describe()` for
            possible aggregation keys. For numeric values, it is recommended to use `["min", "25%", "50%", "75%",
            "max"]` which will result in keeping only the count, mean and std values.
        sort_markdown_columns: If True, the columns of the markdown table are sorted alphabetically.
        markdown_round_digits: The number of digits to round the values in the markdown file to. If None,
            no rounding is applied.
        markdown_data_key: If provided and present in the job return-value, save only the value at this
            key when saving single job results to markdown. This is useful to strip metadata from the job result and,
            thus, allow for correct table formatting of metric results, for instance. If the key is absent (or None),
            a legacy path is used instead, which drops the handle_previous_result field and the result format version
            key from the markdown output.
        multirun_create_ids_from_overrides: If True, create job identifiers from the overrides of the jobs in a
            multi-run. If False, the job index is used as identifier.
        multirun_job_id_key: The key to use for the job identifiers in the integrated multi-run result.
        multirun_convert_job_ids: If True, convert job ids to dictionaries. Works only if
            integrate_multirun_result is True.
        handle_previous_result: If provided, assume the job return-value contains a field with the given
            name (e.g. "prediction") that itself contains an "overrides" field. The overrides from this field are
            either used to replace the existing overrides in the job return object (if replace_existing_overrides is
            True) or simply converted to a dictionary and added back to the job return-value (if
            replace_existing_overrides is False). Furthermore, on the legacy markdown path (see markdown_data_key),
            the field is removed from the job return-value before saving it as markdown to avoid destroying the
            table structure.
        replace_existing_overrides: If True, replace existing overrides in the job return-value with the
            overrides from the job return object if available. If False, the overrides are just converted to a
            dictionary, if available.
        multirun_add_overrides_as_dict: If True, add the overrides as a dictionary to each job return-value
            under the key "overrides".
        multirun_show_file_contents:  A list of filenames (from the filenames attribute or
            aggregated files) whose contents are logged to the console after saving the multi-run results.
            If None is passed, saves [] instead.
        multirun_overrides_separator: The separator to use when creating job identifiers from overrides.
        multirun_markdown_group_by:  The column(s) to group by when saving the multi-run result as
            a markdown file. For numeric columns, the mean and std are calculated. For non-numeric columns, a list of
            values is created. If None, no grouping is applied. A single string is wrapped into a list.
            If a string is passed, it will be wrapped in a list.
        multirun_markdown_transpose: If True, transpose the markdown table for multi-run results.
        paths_file: The file to append the paths of the log directories to. If None, the paths are not
            saved.
        path_id: This is currently not in use. A prefix to add to each line in the paths_file,
            separated by a colon. If None, no prefix is added.
        multirun_paths_file: The file to save the paths of the multi-run log directories to. If None,
            the paths are not saved.
        multirun_path_id: A prefix to add to each line in the multirun_paths_file, separated by a colon.
            If None, no prefix is added.
    """
    self._log = logging.getLogger(f"{__name__}.{self.__class__.__name__}")
    self.filenames = [filenames] if isinstance(filenames, str) else filenames
    self.multirun_show_file_contents = multirun_show_file_contents or []
    self.integrate_multirun_result = integrate_multirun_result
    self._job_returns: list[JobReturn] = []
    self.multirun_aggregator_blacklist = multirun_aggregator_blacklist
    self.sort_markdown_columns = sort_markdown_columns
    self.multirun_create_ids_from_overrides = multirun_create_ids_from_overrides
    self.handle_previous_result = handle_previous_result
    self.replace_existing_overrides = replace_existing_overrides
    self.multirun_add_overrides_as_dict = multirun_add_overrides_as_dict
    self.multirun_job_id_key = multirun_job_id_key
    self.multirun_convert_job_ids = multirun_convert_job_ids
    self.multirun_overrides_separator = multirun_overrides_separator
    if isinstance(multirun_markdown_group_by, str):
        multirun_markdown_group_by = [multirun_markdown_group_by]
    self.multirun_markdown_group_by = multirun_markdown_group_by
    self.multirun_markdown_transpose = multirun_markdown_transpose
    self.markdown_round_digits = markdown_round_digits
    self.markdown_data_key = markdown_data_key
    self.multirun_paths_file = multirun_paths_file
    self.multirun_path_id = multirun_path_id
    self.paths_file = paths_file
    self.path_id = path_id

on_job_end(config, job_return, **kwargs)

Save a single job's return-value to each of the configured filenames via the internal _save method.

Also appends the job to job_returns for later use by on_multirun_end, and appends the job's output directory to paths_file if one is configured. For markdown output, metadata fields are stripped from the return-value first (see markdown_data_key and handle_previous_result) to keep the table structure intact.

Parameters:

Name Type Description Default
config DictConfig

Hydra config of the given job. Only hydra.runtime.output_dir is read.

required
job_return JobReturn

The Hydra job's output object, e.g. the output of predict().

required

Other Parameters:

Name Type Description
**kwargs Any

Ignored; accepted for Hydra callback interface compatibility.

Source code in src/kibad_llm/hydra_callbacks/save_job_return_value.py
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
def on_job_end(self, config: DictConfig, job_return: JobReturn, **kwargs: Any) -> None:
    """Save a single job's return-value to each of the configured filenames via the internal `_save` method.

    Also appends the job to `job_returns` for later use by
    [`on_multirun_end`][kibad_llm.hydra_callbacks.save_job_return_value.SaveJobReturnValueCallback.on_multirun_end],
    and appends the job's output directory to `paths_file` if one is configured. For markdown output, metadata
    fields are stripped from the return-value first (see `markdown_data_key` and `handle_previous_result`) to
    keep the table structure intact.

    Args:
        config: Hydra config of the given job. Only `hydra.runtime.output_dir` is read.
        job_return: The Hydra job's output object, e.g. the output of [`predict()`][kibad_llm.predict.predict].

    Keyword Args:
        **kwargs: Ignored; accepted for Hydra callback interface compatibility.
    """
    if self.handle_previous_result is not None:
        handle_previous_overrides(
            job_return,
            key=self.handle_previous_result,
            replace_existing=self.replace_existing_overrides,
        )
    self._job_returns.append(job_return)
    output_dir = Path(config.hydra.runtime.output_dir)
    if self.paths_file is not None:
        # append the output_dir to the file
        with open(self.paths_file, "a") as file:
            file.write(f"{output_dir}\n")

    for filename in self.filenames:
        # Remove previous result field and "version" (RESULT_FORMAT_VERSION_KEY) from job return-value before
        # saving as markdown. Otherwise, this may destroy the table structure of the saved job return-value.
        obj = job_return.return_value
        if filename.lower().endswith(".md") and isinstance(obj, dict):
            obj = dict(obj)
            if self.markdown_data_key in obj:
                obj = obj[self.markdown_data_key]
            else:  # legacy code path for compatibility
                if (
                    self.handle_previous_result is not None
                    and self.handle_previous_result in obj
                ):
                    obj.pop(self.handle_previous_result)
                if RESULT_FORMAT_VERSION_KEY in obj:
                    obj.pop(RESULT_FORMAT_VERSION_KEY)
                if self.markdown_data_key in obj:
                    obj = obj[self.markdown_data_key]
        self._save(obj=obj, filename=filename, output_dir=output_dir)

on_multirun_end(config, **kwargs)

Collate a multi-run and all its jobs' data, then save it.

Parameters:

Name Type Description Default
config DictConfig

The multi-run's Hydra config.

required

Other Parameters:

Name Type Description
**kwargs Any

Ignored; accepted for Hydra callback interface compatibility.

Source code in src/kibad_llm/hydra_callbacks/save_job_return_value.py
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
def on_multirun_end(self, config: DictConfig, **kwargs: Any) -> None:
    """Collate a multi-run and all its jobs' data, then save it.

    Args:
        config: The multi-run's Hydra config.

    Keyword Args:
        **kwargs: Ignored; accepted for Hydra callback interface compatibility.
    """
    job_ids: list[str] | list[int] | None = None
    if self.multirun_create_ids_from_overrides:
        overrides_per_result = [jr.overrides or [] for jr in self._job_returns]
        job_ids = overrides_to_identifiers(
            overrides_per_result, sep=self.multirun_overrides_separator, remove_common=True
        )
        if job_ids is None:
            self._log.warning(
                "Job identifiers created from overrides are not unique! "
                "Use the job indexes instead."
            )

    if job_ids is None:
        job_ids = list[int](range(len(self._job_returns)))

    if self.multirun_add_overrides_as_dict:
        for jr in self._job_returns:
            jr.return_value["overrides"] = overrides_to_dict(
                jr.overrides or [], remove_plus_prefix=True
            )

    if self.integrate_multirun_result:
        # WARN: list_of_dicts may return lists. There is a safety backup (the {"value": obj} wrapper),
        #   but this is very sketchy.
        #
        # rearrange the job return-values of all jobs from a multi-run into a dict of lists (maybe nested),
        obj = list_of_dicts_to_dict_of_lists_recursive(
            [jr.return_value for jr in self._job_returns]
        )
        if not isinstance(obj, dict):
            obj = {"value": obj}
        if self.multirun_create_ids_from_overrides:
            obj[self.multirun_job_id_key] = job_ids

        # also create an aggregated result
        # convert to python object to allow selecting numeric columns
        obj_py = to_py_obj(obj)
        obj_flat = flatten_dict(cast(dict[str | int, Any], obj_py))
        # create dataframe from flattened dict
        df_flat = pd.DataFrame(obj_flat)
        # select only the numeric values
        df_numbers_only = df_flat.select_dtypes(["number"])
        cols_removed: set[tuple[str | int, ...]] = (
            set(df_flat.columns) - set(df_numbers_only.columns) - {(self.multirun_job_id_key,)}  # type: ignore
        )
        if len(cols_removed) > 0:
            self._log.warning(
                f"Removed the following columns from the aggregated result because they are not numeric: "
                f"{cols_removed}"
            )
        if len(df_numbers_only.columns) == 0:
            obj_aggregated = None
        else:
            # aggregate the numeric values
            df_described = df_numbers_only.describe()
            # remove rows in the blacklist
            if self.multirun_aggregator_blacklist is not None:
                df_described = df_described.drop(
                    self.multirun_aggregator_blacklist, errors="ignore", axis="index"
                )
            # add the aggregation keys (e.g. mean, min, ...) as most inner keys and convert back to dict
            # TODO: check if "type ignore" is really fine and necessary here
            obj_flat_aggregated: dict[tuple[str | int | float, ...], Any] = df_described.T.stack().to_dict()  # type: ignore
            # unflatten because _save() works better with nested dicts. But don't remove key padding
            # since this is required for proper unstacking in _save() for markdown files.
            obj_aggregated = unflatten_dict(obj_flat_aggregated, unpad_keys=False)

        if self.multirun_convert_job_ids:
            # convert job ids (created from overrides) to dicts
            obj[self.multirun_job_id_key] = list_of_dicts_to_dict_of_lists_recursive(
                [
                    identifier_to_dict(identifier, sep=self.multirun_overrides_separator)
                    for identifier in obj[self.multirun_job_id_key]
                ]
            )
    else:
        # create a dict of the job return-values of all jobs from a multi-run
        # (_save() works better with nested dicts)
        obj = {
            identifier: jr.return_value for identifier, jr in zip(job_ids, self._job_returns)
        }
        obj_aggregated = None
    output_dir = Path(config.hydra.sweep.dir)
    if self.multirun_paths_file is not None:
        # append the output_dir to the file
        line = f"{output_dir}\n"
        if self.multirun_path_id is not None:
            line = f"{self.multirun_path_id}:{line}"
        with open(self.multirun_paths_file, "a") as file:
            file.write(line)

    filenames_aggregated = []
    for filename in self.filenames:
        self._save(
            obj=obj,
            filename=filename,
            output_dir=output_dir,
            is_tabular_data=self.integrate_multirun_result,
            markdown_group_by=self.multirun_markdown_group_by,
        )
        # if available, also save the aggregated result
        if obj_aggregated is not None:
            file_base_name, ext = os.path.splitext(filename)
            filename_aggregated = f"{file_base_name}.aggregated{ext}"
            filenames_aggregated.append(filename_aggregated)
            self._save(
                obj=obj_aggregated,
                filename=filename_aggregated,
                output_dir=output_dir,
                # If we have aggregated (integrated multi-run) results, we unstack the last level,
                # i.e. the aggregation key.
                unstack_last_index_level=True,
            )
    saved_files = set(self.filenames + filenames_aggregated)
    for fn in self.multirun_show_file_contents:
        if fn in saved_files:
            with open(str(output_dir / fn)) as file:
                contents = file.read()
            self._log.info(f"Contents of {output_dir / fn}:\n{contents}")