-
Enhancing Retrieval for ESGLLM via ESG-CID -- A Disclosure Content Index Finetuning Dataset for Mapping GRI and ESRS
Paper • 2503.10674 • Published • 3 -
ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge
Paper • 2506.01646 • Published • 2 -
Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs
Paper • 2406.19102 • Published • 1 -
Improving Retrieval for RAG based Question Answering Models on Financial Documents
Paper • 2404.07221 • Published
Guille Pérez-Torró
guishe
AI & ML interests
Information Retrieval, Few-Shot Learning, Named Entity Recognition, Named Entity Disambiguation, Semantic Search, Aspect-based Sentiment Analysis
Recent Activity
upvoted a paper about 2 months ago
Autodata: An agentic data scientist to create high quality synthetic data liked a Space 2 months ago
HuggingFaceH4/on-policy-distillation liked a Space 2 months ago
HuggingFaceTB/trl-distillation-trainerOrganizations
Multimodal Embeddings
Reasoning LLMs
ReRanker Encoder-only Models
-
jinaai/jina-reranker-v2-base-multilingual
Text Ranking • 0.3B • Updated • 1.15M • 355 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 18.9M • • 1.14k -
mixedbread-ai/mxbai-rerank-base-v1
Text Ranking • 0.2B • Updated • 48.1k • 46 -
mixedbread-ai/mxbai-rerank-large-v1
Text Ranking • 0.4B • Updated • 29.5k • 143
Embedding Encoder-only Models
-
BAAI/bge-m3
Sentence Similarity • Updated • 35.6M • • 3.4k -
BAAI/bge-large-en-v1.5
Feature Extraction • 0.3B • Updated • 12.7M • • 714 -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 16.3M • 892 -
mixedbread-ai/mxbai-embed-large-v1
Feature Extraction • 0.3B • Updated • 4.16M • • 822
Summarization
-
Falconsai/text_summarization
Summarization • 60.5M • Updated • 43.3k • • 303 -
knkarthick/MEETING-SUMMARY-BART-LARGE-XSUM-SAMSUM-DIALOGSUM
Summarization • 0.4B • Updated • 143 • 13 -
knkarthick/MEETING-SUMMARY-BART-LARGE-XSUM-SAMSUM-DIALOGSUM-AMI
Summarization • 0.4B • Updated • 33 • 17 -
BEE-spoke-data/pegasus-x-base-synthsumm_open-16k
Summarization • 0.3B • Updated • 85 • 3
Instruct LLMs
-
instruction-pretrain/finance-Llama3-8B
Text Generation • 8B • Updated • 13.5k • 80 -
EmergentMethods/Phi-3-mini-4k-instruct-graph
Text Generation • 4B • Updated • 66 • 48 -
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
Text Generation • 71B • Updated • 10.6k • • 2.07k -
unsloth/Llama-3.2-3B-Instruct-unsloth-bnb-4bit
Text Generation • 3B • Updated • 93.1k • 10
Multi-Vector Embedding Models
LLM-as-a-Judge
-
AtlaAI/Selene-1-Mini-Llama-3.1-8B
Text Generation • 8B • Updated • 1.69k • • 107 -
AtlaAI/Selene-1-Mini-Llama-3.1-8B-Q4_K_M-GGUF
Text Generation • 8B • Updated • 159 • 8 -
flowaicom/Flow-Judge-v0.1
Text Generation • 4B • Updated • 4.48k • 71 -
prometheus-eval/prometheus-7b-v2.0
Text Generation • 7B • Updated • 13.7k • • 109
Zero-Shot Entailment Models
-
MoritzLaurer/bge-m3-zeroshot-v2.0
Zero-Shot Classification • 0.6B • Updated • 119k • • 68 -
MoritzLaurer/deberta-v3-large-zeroshot-v2.0
Zero-Shot Classification • 0.4B • Updated • 122k • • 133 -
MoritzLaurer/ModernBERT-base-zeroshot-v2.0
Text Classification • 0.1B • Updated • 1.11k • • 19 -
MoritzLaurer/ModernBERT-large-zeroshot-v2.0
Text Classification • 0.4B • Updated • 14.1k • • 69
NER Encoder-only Models
This collections gathers several NER models. Either fine-tuned versions for specific tasks or generic backbone models ready to be fine-tuned.
-
guishe/nuner-v1_orgs
Token Classification • 0.1B • Updated • 815 • 2 -
guishe/span-marker-generic-ner-v1-fewnerd-fine-super
Token Classification • 0.1B • Updated • 2.98k • 13 -
protectai/guishe-nuner-v1_orgs-onnx
Token Classification • Updated • 371 -
guishe/nuner-v1_fewnerd_fine_super
Token Classification • 0.1B • Updated • 7
Small LLMs
-
microsoft/Phi-3-mini-4k-instruct
Text Generation • 4B • Updated • 614k • • 1.45k -
google/gemma-2-2b-it
Text Generation • 3B • Updated • 741k • • 1.45k -
nvidia/Nemotron-Mini-4B-Instruct
Text Generation • Updated • 9.89k • 186 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 433k • 211
Part-of-Speech Tagging
papers_ESG
-
Enhancing Retrieval for ESGLLM via ESG-CID -- A Disclosure Content Index Finetuning Dataset for Mapping GRI and ESRS
Paper • 2503.10674 • Published • 3 -
ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge
Paper • 2506.01646 • Published • 2 -
Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs
Paper • 2406.19102 • Published • 1 -
Improving Retrieval for RAG based Question Answering Models on Financial Documents
Paper • 2404.07221 • Published
Multi-Vector Embedding Models
Multimodal Embeddings
LLM-as-a-Judge
-
AtlaAI/Selene-1-Mini-Llama-3.1-8B
Text Generation • 8B • Updated • 1.69k • • 107 -
AtlaAI/Selene-1-Mini-Llama-3.1-8B-Q4_K_M-GGUF
Text Generation • 8B • Updated • 159 • 8 -
flowaicom/Flow-Judge-v0.1
Text Generation • 4B • Updated • 4.48k • 71 -
prometheus-eval/prometheus-7b-v2.0
Text Generation • 7B • Updated • 13.7k • • 109
Reasoning LLMs
Zero-Shot Entailment Models
-
MoritzLaurer/bge-m3-zeroshot-v2.0
Zero-Shot Classification • 0.6B • Updated • 119k • • 68 -
MoritzLaurer/deberta-v3-large-zeroshot-v2.0
Zero-Shot Classification • 0.4B • Updated • 122k • • 133 -
MoritzLaurer/ModernBERT-base-zeroshot-v2.0
Text Classification • 0.1B • Updated • 1.11k • • 19 -
MoritzLaurer/ModernBERT-large-zeroshot-v2.0
Text Classification • 0.4B • Updated • 14.1k • • 69
ReRanker Encoder-only Models
-
jinaai/jina-reranker-v2-base-multilingual
Text Ranking • 0.3B • Updated • 1.15M • 355 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 18.9M • • 1.14k -
mixedbread-ai/mxbai-rerank-base-v1
Text Ranking • 0.2B • Updated • 48.1k • 46 -
mixedbread-ai/mxbai-rerank-large-v1
Text Ranking • 0.4B • Updated • 29.5k • 143
NER Encoder-only Models
This collections gathers several NER models. Either fine-tuned versions for specific tasks or generic backbone models ready to be fine-tuned.
-
guishe/nuner-v1_orgs
Token Classification • 0.1B • Updated • 815 • 2 -
guishe/span-marker-generic-ner-v1-fewnerd-fine-super
Token Classification • 0.1B • Updated • 2.98k • 13 -
protectai/guishe-nuner-v1_orgs-onnx
Token Classification • Updated • 371 -
guishe/nuner-v1_fewnerd_fine_super
Token Classification • 0.1B • Updated • 7
Embedding Encoder-only Models
-
BAAI/bge-m3
Sentence Similarity • Updated • 35.6M • • 3.4k -
BAAI/bge-large-en-v1.5
Feature Extraction • 0.3B • Updated • 12.7M • • 714 -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 16.3M • 892 -
mixedbread-ai/mxbai-embed-large-v1
Feature Extraction • 0.3B • Updated • 4.16M • • 822
Small LLMs
-
microsoft/Phi-3-mini-4k-instruct
Text Generation • 4B • Updated • 614k • • 1.45k -
google/gemma-2-2b-it
Text Generation • 3B • Updated • 741k • • 1.45k -
nvidia/Nemotron-Mini-4B-Instruct
Text Generation • Updated • 9.89k • 186 -
HuggingFaceTB/SmolLM2-360M-Instruct
Text Generation • 0.4B • Updated • 433k • 211
Summarization
-
Falconsai/text_summarization
Summarization • 60.5M • Updated • 43.3k • • 303 -
knkarthick/MEETING-SUMMARY-BART-LARGE-XSUM-SAMSUM-DIALOGSUM
Summarization • 0.4B • Updated • 143 • 13 -
knkarthick/MEETING-SUMMARY-BART-LARGE-XSUM-SAMSUM-DIALOGSUM-AMI
Summarization • 0.4B • Updated • 33 • 17 -
BEE-spoke-data/pegasus-x-base-synthsumm_open-16k
Summarization • 0.3B • Updated • 85 • 3
Part-of-Speech Tagging
Instruct LLMs
-
instruction-pretrain/finance-Llama3-8B
Text Generation • 8B • Updated • 13.5k • 80 -
EmergentMethods/Phi-3-mini-4k-instruct-graph
Text Generation • 4B • Updated • 66 • 48 -
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
Text Generation • 71B • Updated • 10.6k • • 2.07k -
unsloth/Llama-3.2-3B-Instruct-unsloth-bnb-4bit
Text Generation • 3B • Updated • 93.1k • 10