scitldr

参考文献:

抽象的な

次のコマンドを使用して、このデータセットを TFDS にロードします。

ds = tfds.load('huggingface:scitldr/Abstract')
  • 説明
A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
SCITLDR contains both author-written and expert-derived TLDRs,
where the latter are collected using a novel annotation protocol
that produces high-quality summaries while minimizing annotation burden.
  • ライセンス: Apache ライセンス 2.0
  • バージョン: 0.0.0
  • 分割:
スプリット
'test' 618
'train' 1992年
'validation' 619
  • 特徴
{
    "source": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "source_labels": {
        "feature": {
            "num_classes": 2,
            "names": [
                "non-oracle",
                "oracle"
            ],
            "names_file": null,
            "id": null,
            "_type": "ClassLabel"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "rouge_scores": {
        "feature": {
            "dtype": "float32",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "paper_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "target": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    }
}

AIC

次のコマンドを使用して、このデータセットを TFDS にロードします。

ds = tfds.load('huggingface:scitldr/AIC')
  • 説明
A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
SCITLDR contains both author-written and expert-derived TLDRs,
where the latter are collected using a novel annotation protocol
that produces high-quality summaries while minimizing annotation burden.
  • ライセンス: Apache ライセンス 2.0
  • バージョン: 0.0.0
  • 分割:
スプリット
'test' 618
'train' 1992年
'validation' 619
  • 特徴
{
    "source": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "source_labels": {
        "feature": {
            "num_classes": 2,
            "names": [
                0,
                1
            ],
            "names_file": null,
            "id": null,
            "_type": "ClassLabel"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "rouge_scores": {
        "feature": {
            "dtype": "float32",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "paper_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "ic": {
        "dtype": "bool_",
        "id": null,
        "_type": "Value"
    },
    "target": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    }
}

全文

次のコマンドを使用して、このデータセットを TFDS にロードします。

ds = tfds.load('huggingface:scitldr/FullText')
  • 説明
A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
SCITLDR contains both author-written and expert-derived TLDRs,
where the latter are collected using a novel annotation protocol
that produces high-quality summaries while minimizing annotation burden.
  • ライセンス: Apache ライセンス 2.0
  • バージョン: 0.0.0
  • 分割:
スプリット
'test' 618
'train' 1992年
'validation' 619
  • 特徴
{
    "source": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "source_labels": {
        "feature": {
            "num_classes": 2,
            "names": [
                "non-oracle",
                "oracle"
            ],
            "names_file": null,
            "id": null,
            "_type": "ClassLabel"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "rouge_scores": {
        "feature": {
            "dtype": "float32",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    },
    "paper_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "target": {
        "feature": {
            "dtype": "string",
            "id": null,
            "_type": "Value"
        },
        "length": -1,
        "id": null,
        "_type": "Sequence"
    }
}