Firecrawl Alexandria
Hugging Face
Codefollows the fields above
const result = await firecrawl.scrape({
alexandria: {
provider: "huggingface-co",
capability: "hub/card",
options: {
repo_id: "google-bert/bert-base-uncased",
},
},
});2Response example
This is a sample shape. Press Run to see live data from Hugging Face.
Long values are shortened in this preview. The full response is 12 KB.
{
"body": "# BERT base model (uncased)\n\nPretrained model on English language using a masked language modeling (MLM) objective. It was introduced in\n[this paper](https://arxiv.org/abs/1810.04805) and first released in\n[this repository](https://github.com/google-research/bert). This model is uncased: it does not make a difference\nbetween english and English.\n\nDisclaimer: The team releasing BERT did not write a model card for this model so this model card has been written by\nthe Hugging Face team.\n\n## Model description\n\nBERT is a transformers model pretrained on a large corpus of English data in a self-supervised fashion. This means it\nwas pretrained on the raw texts only, with no humans labeling them in any way (which is why it can use lots of\npublicly available data) with an automatic process to generate inputs and labels from those texts. More precisely, it\nwas pretrained with two objectives:\n\n- Masked language modeling (MLM): taking a sentence, the model randomly masks 15% of the words in the input then run\n the entire masked sentence through the model and has to predict the masked words. This is different from traditional\n recurrent neural networks (RNNs) that usually see the words one after the other, or from autoregressive models like\n GPT which internally masks the future tokens. It allows the model to learn a bidirectional representation of the\n sentence.\n- Next sentence prediction (NSP): the models concatenates two masked sentences as inputs during pretraining. Sometimes\n they correspond to sentences that were next to each other in the original text, sometimes not. The model then has to\n predict if the two sentences were following each other or not.\n\nThis way, the model learns an inner representation of the English language that can then be used to extract features\nuseful for downstream tasks: if you have a dataset of labeled sentences, for instance, you can train a standard\nclassifier using the features produced by the BERT model as inputs.\n\n## Model variations\n\n… [8,425 more characters]",
"card_data": {
"datasets": [
"bookcorpus",
"wikipedia"
],
"language": "en",
"license": "apache-2.0",
"tags": [
"exbert"
]
},
"front_matter": "language: en\ntags:\n- exbert\nlicense: apache-2.0\ndatasets:\n- bookcorpus\n- wikipedia",
"gated": false,
"gated_mode": null,
"has_card": true,
"last_modified": "2024-02-19T11:06:12.000Z",
"license": "apache-2.0",
"observed_at_ms": 1789431352194,
"repo_id": "google-bert/bert-base-uncased",
"repo_type": "model",
"requested_repo_id": "google-bert/bert-base-uncased",
"revision": "main",
"sections": [
{
"level": 1,
"title": "BERT base model (uncased)"
},
{
"level": 2,
"title": "Model description"
},
{
"level": 2,
"title": "Model variations"
},
{
"level": 2,
"title": "Intended uses & limitations"
},
{
"level": 3,
"title": "How to use"
},
{
"level": 3,
"title": "Limitations and bias"
},
{
"level": 2,
"title": "Training data"
},
{
"level": 2,
"title": "Training procedure"
},
{
"level": 3,
"title": "Preprocessing"
},
{
"level": 3,
"title": "Pretraining"
},
{
"level": 2,
"title": "Evaluation results"
},
{
"level": 3,
"title": "BibTeX entry and citation info"
}
],
"source_url": "https://huggingface.co/google-bert/bert-base-uncased/resolve/main/README.md",
"url": "https://huggingface.co/google-bert/bert-base-uncased",
"word_count": 1357
}