# F-UJI

[F-UJI](https://github.com/pangaea-data-publisher/fuji) checks dataset metadata
against FAIR requirements.

## Supported versions

| F-UJI version   | Metric version | Total metrics | Metrics supported offline | Total checks | Checks supported offline |
| --------------- | -------------- | ------------- | ------------------------- | ------------ | ------------------------ |
| 3.5.1 (default) | 0.8            | 17            | 13 full + 1 partial       | 31           | 24                       |

Full metric support means all its checks are supported; partial means only some.
A supported check still needs the required information to produce a decision.

Select a version explicitly to keep using its mappings, checks and scoring rules:

```python
from fair_offline_assessor import Assessor

assessor = Assessor("FUJI", version="3.5.1")
```

## What it checks

F-UJI looks for identifiers, descriptive information, licences, access conditions,
links to data, relationships between datasets, and information about how data was
created. Some checks examine whether recognised standards and file formats are used.

Each check has its own requirements and points. These examples use F-UJI 3.5.1:

| Required information                                                     | Check ID         | Points when met |
| ------------------------------------------------------------------------ | ---------------- | --------------- |
| `metadata_url` with recognised identifier syntax, such as a URL or UUID  | `FsF-F1-01MD-1`  | 1               |
| Creator, title, identifier, publication date, publisher and dataset type | `FsF-F2-01M-2`   | 1               |
| All of the preceding descriptive fields, plus a summary and keywords     | `FsF-F2-01M-3`   | 1               |
| Licence information in an appropriate metadata field                     | `FsF-R1.1-01M-1` | 1               |

Field names depend on the metadata standard you use. F-UJI's mappings connect
those fields to the information each check requires. For example, Schema.org
`name` supplies the dataset title.

The complete requirements and points are in
[F-UJI's check definitions for this version](https://github.com/pangaea-data-publisher/fuji/blob/9227fabb7f047475714f2e7622798b855c883f72/fuji_server/yaml/metrics_v0.8.yaml).

## What can run offline

The seven unsupported checks remain `indeterminate` with reason `unsupported_check`.

Supported checks may also remain indeterminate if the required information is
missing or cannot be interpreted. Offline results cannot confirm that a link
opens, a dataset is downloadable or a search engine can find it.

## Metadata

Pass document contents to `metadata`. JSON-LD is the default. For another format,
set `metadata_format` to one of the values below.

| `metadata_format`                                | Accepted contents                                                                             |
| ------------------------------------------------ | --------------------------------------------------------------------------------------------- |
| omitted or `json-ld`                             | A JSON-LD dictionary, list of dictionaries or JSON text                                       |
| `datacite-json`                                  | DataCite JSON with `agency` and metadata fields such as `titles` at the top level             |
| `xml`                                            | DataCite, Dublin Core, DDI, CMDI, DIF, MODS, EML, ISO metadata, EAD or TEI XML                |
| `html`                                           | HTML containing JSON-LD, Dublin Core, Microdata, RDFa, Highwire/Eprints or OpenGraph metadata |
| `turtle`, `n3`, `nt`, `nquads`, `trig`, `rdfxml` | Text in the selected RDF format                                                               |

DataCite JSON requires `agency`. DataCite API responses with metadata inside
`data.attributes` are not supported. XML must contain its own information, without
DOCTYPE or entity declarations. If an OAI-PMH XML response contains several
records, F-UJI uses the first. RSS/GeoRSS and OAI-ORE Atom are not supported.

Supply one intended dataset per document. F-UJI decides which described dataset
to assess using its own rules. With JSON-LD or other RDF formats, you can set
`subject` to the identifier you expect it to choose. A choice that cannot be
confirmed produces `subject_not_selected`. Omit `subject` for XML, DataCite JSON
and HTML.

For Schema.org metadata:

* Use `identifier` for the dataset identifier required by the citation check.
  F-UJI does not use `@id` for that field.
* Describe creators as `Person` or `Organization` objects, for example
  `"creator": {"@type": "Person", "name": "Ada Example"}`. F-UJI 3.5.1 does not
  use a plain text `creator` value in RDF input.
* `isAccessibleForFree: false` is not used by this version's RDF mapping.
  That value alone cannot establish access restrictions.

**DCAT metadata with Dublin Core properties**

```python
from fair_offline_assessor import Assessor

result = Assessor("FUJI", version="3.5.1").assess(
    metadata={
        "@context": {
            "dcat": "http://www.w3.org/ns/dcat#",
            "dct": "http://purl.org/dc/terms/",
        },
        "@type": "dcat:Dataset",
        "@id": "https://example.org/datasets/1",
        "dct:title": "Example dataset",
        "dct:license": {
            "@id": "https://creativecommons.org/licenses/by/4.0/"
        },
    },
    metadata_url="https://example.org/datasets/1",
)
```

Schema.org, DCAT and Dublin Core are **vocabularies**: sets of named properties
with agreed meanings. DataCite defines a **metadata schema**: fields and rules
for describing a dataset. **RDF** describes things through statements such as
“this dataset has this title”. JSON-LD and Turtle are ways to write those statements.

## Contexts

A JSON-LD `@context` defines what field names mean. The library includes the
Schema.org context, so `"@context": "https://schema.org"` works offline.

For another context URL, supply its document in `local_contexts`, or put the
field definitions directly inside `@context`. Any other contexts it refers to
must also be supplied. Context definitions cannot override different definitions
already included in the library.

**Supply the definitions for a context URL**

```python
from fair_offline_assessor import Assessor

result = Assessor("FUJI", version="3.5.1").assess(
    metadata={
        "@context": "https://example.org/context",
        "@type": "Dataset",
        "name": "Example dataset",
    },
    local_contexts={
        "https://example.org/context": {
            "@context": {
                "@vocab": "https://schema.org/"
            }
        }
    },
)
```

## Results

The common response includes scores for individual checks and metrics.
The `evidence` entries refer to your whole metadata document or `metadata_url`.

`raw` contains unmodified F-UJI result dictionaries, with at most one entry per
metric. Metrics that produce no result are omitted. It covers this offline run,
not a full response from F-UJI's online service.

See [Results](/local-offline-assessor-for-fair-loaf/results) for field meanings and
[Troubleshooting](#troubleshooting) for problems with metadata.

## Troubleshooting

Start with `result.diagnostics` for problems with the metadata. For a particular
check, read its `reason_code` and `message`.

`indeterminate` means the library cannot decide whether the check passes.
Other checks can still run if they have the information they need.

### Example: a missing context

A JSON-LD context defines what field names mean. In this example, the metadata
refers to a context URL whose definitions are not available to the library.
It returns `unknown_context` because it cannot download them.

Supply the document through `local_contexts`, or put the definitions directly
inside `@context`. See [F-UJI contexts](/local-offline-assessor-for-fair-loaf/assessors/fuji#contexts).

The assessment can still have `status="completed"`. That means processing
finished; the diagnostics explain which parts could not be assessed.

**A context that was not supplied**

```python
from fair_offline_assessor import Assessor

result = Assessor("FUJI", version="3.5.1").assess(
    metadata={
        "@context": "https://example.org/context",
        "@type": "Dataset",
    }
)

print(result.diagnostics[0].model_dump_json(indent=2))
```

**Returned message**

```json
{
  "code": "unknown_context",
  "message": "Context unavailable locally: https://example.org/context",
  "location": "/metadata"
}
```

### Problems with metadata

| Code                          | Meaning or action                                                                                                                |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| `invalid_json`                | Correct the JSON syntax. Values such as `NaN` are not valid JSON numbers.                                                        |
| `unsupported_input`           | Supply a dictionary, a list of dictionaries or document text as required by the chosen format.                                   |
| `invalid_jsonld`              | Correct the JSON-LD problem identified in the message.                                                                           |
| `unknown_context`             | Supply the definitions for the context URL.                                                                                      |
| `invalid_context`             | Include `@context` in the supplied context document.                                                                             |
| `context_conflict`            | The supplied context differs from one included in the library. Use the included definitions or a different context URL.          |
| `metadata_reader_error`       | F-UJI could not interpret the metadata. Check that its structure matches the chosen format.                                      |
| `invalid_rdf`                 | Correct the RDF syntax identified in the message.                                                                                |
| `invalid_embedded_metadata`   | Correct the metadata inside the HTML. Other usable information is still assessed.                                                |
| `invalid_rdfa`                | Correct the RDFa metadata inside the HTML. Other usable information is still assessed.                                           |
| `metadata_not_found`          | F-UJI found no usable metadata. Check the contents and chosen format.                                                            |
| `unsupported_metadata_format` | Choose one of the supported `metadata_format` values.                                                                            |
| `unsupported_datacite_shape`  | Supply DataCite JSON with `agency` and its metadata fields at the top level. Metadata inside `data.attributes` is not supported. |
| `subject_not_selected`        | F-UJI could not confirm the dataset specified by `subject`. Supply a document describing the intended dataset.                   |
| `unsupported_subject`         | Omit `subject` for XML, DataCite JSON and HTML.                                                                                  |
| `unsafe_xml`                  | Remove DOCTYPE and entity declarations. Supply XML that contains its own information.                                            |
| `offline_reference`           | The assessment cannot fetch the referenced information.                                                                          |

Codes such as `invalid_subject`, `invalid_metadata_url` and `invalid_local_contexts`
mean an argument has the wrong type or value. Check the message and the
[Python API](/local-offline-assessor-for-fair-loaf/reference#assess-a-dataset).

### Why a check could not run

| Check reason        | Meaning                                                                                                                    |
| ------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| `missing_evidence`  | The information needed to decide this check was not supplied.                                                              |
| `unsupported_check` | The library cannot run this check offline.                                                                                 |
| `evaluator_error`   | An error prevented the check from finishing. Its outcome is `error`, and the assessment status is `completed_with_errors`. |

A metadata problem's code can also appear as a check's reason when that problem
prevents the check from running. To report an error, include the library version,
check ID and a small example that reproduces it.

### Choosing an assessor

`Assessor(...)` raises `ProfileError` if the requested assessor or version is
unavailable. Its `code` is `assessor_not_found` or `assessor_version_not_found`.
Use `Assessor("FUJI")` to select the current default.
