opus_wikipedia

مراجع:

ar-en

برای بارگذاری این مجموعه داده در TFDS از دستور زیر استفاده کنید:

ds = tfds.load('huggingface:opus_wikipedia/ar-en')

توضیحات :

This is a corpus of parallel sentences extracted from Wikipedia by Krzysztof Wołk and Krzysztof Marasek. Please cite the following publication if you use the data: Krzysztof Wołk and Krzysztof Marasek: Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs., Procedia Technology, 18, Elsevier, p.126-132, 2014
20 languages, 36 bitexts
total number of files: 114
total number of tokens: 610.13M
total number of sentence fragments: 25.90M

مجوز : مجوز شناخته شده ای وجود ندارد
نسخه : 1.0.0
تقسیمات :

تقسیم کنید	نمونه ها
`'train'`	151136

ویژگی ها :

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "ar",
            "en"
        ],
        "id": null,
        "_type": "Translation"
    }
}

ar-pl

برای بارگذاری این مجموعه داده در TFDS از دستور زیر استفاده کنید:

ds = tfds.load('huggingface:opus_wikipedia/ar-pl')

توضیحات :

This is a corpus of parallel sentences extracted from Wikipedia by Krzysztof Wołk and Krzysztof Marasek. Please cite the following publication if you use the data: Krzysztof Wołk and Krzysztof Marasek: Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs., Procedia Technology, 18, Elsevier, p.126-132, 2014
20 languages, 36 bitexts
total number of files: 114
total number of tokens: 610.13M
total number of sentence fragments: 25.90M

مجوز : مجوز شناخته شده ای وجود ندارد
نسخه : 1.0.0
تقسیمات :

تقسیم کنید	نمونه ها
`'train'`	823715

ویژگی ها :

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "ar",
            "pl"
        ],
        "id": null,
        "_type": "Translation"
    }
}

en-sl

برای بارگذاری این مجموعه داده در TFDS از دستور زیر استفاده کنید:

ds = tfds.load('huggingface:opus_wikipedia/en-sl')

توضیحات :

This is a corpus of parallel sentences extracted from Wikipedia by Krzysztof Wołk and Krzysztof Marasek. Please cite the following publication if you use the data: Krzysztof Wołk and Krzysztof Marasek: Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs., Procedia Technology, 18, Elsevier, p.126-132, 2014
20 languages, 36 bitexts
total number of files: 114
total number of tokens: 610.13M
total number of sentence fragments: 25.90M

مجوز : مجوز شناخته شده ای وجود ندارد
نسخه : 1.0.0
تقسیمات :

تقسیم کنید	نمونه ها
`'train'`	140124

ویژگی ها :

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "en",
            "sl"
        ],
        "id": null,
        "_type": "Translation"
    }
}

en-ru

برای بارگذاری این مجموعه داده در TFDS از دستور زیر استفاده کنید:

ds = tfds.load('huggingface:opus_wikipedia/en-ru')

توضیحات :

This is a corpus of parallel sentences extracted from Wikipedia by Krzysztof Wołk and Krzysztof Marasek. Please cite the following publication if you use the data: Krzysztof Wołk and Krzysztof Marasek: Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs., Procedia Technology, 18, Elsevier, p.126-132, 2014
20 languages, 36 bitexts
total number of files: 114
total number of tokens: 610.13M
total number of sentence fragments: 25.90M

مجوز : مجوز شناخته شده ای وجود ندارد
نسخه : 1.0.0
تقسیمات :

تقسیم کنید	نمونه ها
`'train'`	572717

ویژگی ها :

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "en",
            "ru"
        ],
        "id": null,
        "_type": "Translation"
    }
}

en-vi

برای بارگذاری این مجموعه داده در TFDS از دستور زیر استفاده کنید:

ds = tfds.load('huggingface:opus_wikipedia/en-vi')

توضیحات :

This is a corpus of parallel sentences extracted from Wikipedia by Krzysztof Wołk and Krzysztof Marasek. Please cite the following publication if you use the data: Krzysztof Wołk and Krzysztof Marasek: Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs., Procedia Technology, 18, Elsevier, p.126-132, 2014
20 languages, 36 bitexts
total number of files: 114
total number of tokens: 610.13M
total number of sentence fragments: 25.90M

مجوز : مجوز شناخته شده ای وجود ندارد
نسخه : 1.0.0
تقسیمات :

تقسیم کنید	نمونه ها
`'train'`	58116

ویژگی ها :

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "en",
            "vi"
        ],
        "id": null,
        "_type": "Translation"
    }
}