scitldr

参考文献:

抽象的な

次のコマンドを使用して、このデータセットを TFDS にロードします。

ds = tfds.load('huggingface:scitldr/Abstract')

説明：

A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
SCITLDR contains both author-written and expert-derived TLDRs,
where the latter are collected using a novel annotation protocol
that produces high-quality summaries while minimizing annotation burden.

ライセンス: Apache ライセンス 2.0
バージョン: 0.0.0
分割:

スプリット	例
`'test'`	618
`'train'`	1992年
`'validation'`	619

特徴：

{
    "source": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "source_labels": {
        "feature": {
            "num_classes": 2,
            "names": [
                "non-oracle",
                "oracle"
            ],
            "names_file": null,
            "id": null,
            "_type": "ClassLabel"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "rouge_scores": {
        "feature": {
            "dtype": "float32",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "paper_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "target": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    }
}

AIC

次のコマンドを使用して、このデータセットを TFDS にロードします。

ds = tfds.load('huggingface:scitldr/AIC')

説明：

A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
SCITLDR contains both author-written and expert-derived TLDRs,
where the latter are collected using a novel annotation protocol
that produces high-quality summaries while minimizing annotation burden.

ライセンス: Apache ライセンス 2.0
バージョン: 0.0.0
分割:

スプリット	例
`'test'`	618
`'train'`	1992年
`'validation'`	619

特徴：

{
    "source": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "source_labels": {
        "feature": {
            "num_classes": 2,
            "names": [
                0,
                1
            ],
            "names_file": null,
            "id": null,
            "_type": "ClassLabel"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "rouge_scores": {
        "feature": {
            "dtype": "float32",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "paper_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "ic": {
        "dtype": "bool_",
        "id": null,
        "_type": "Value"
    },
    "target": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    }
}

全文

次のコマンドを使用して、このデータセットを TFDS にロードします。

ds = tfds.load('huggingface:scitldr/FullText')

説明：

A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
SCITLDR contains both author-written and expert-derived TLDRs,
where the latter are collected using a novel annotation protocol
that produces high-quality summaries while minimizing annotation burden.

ライセンス: Apache ライセンス 2.0
バージョン: 0.0.0
分割:

スプリット	例
`'test'`	618
`'train'`	1992年
`'validation'`	619

特徴：

{
    "source": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "source_labels": {
        "feature": {
            "num_classes": 2,
            "names": [
                "non-oracle",
                "oracle"
            ],
            "names_file": null,
            "id": null,
            "_type": "ClassLabel"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "rouge_scores": {
        "feature": {
            "dtype": "float32",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "paper_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "target": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    }
}