Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster
With the development of large language models (LLMs), generative retrieval is becoming increasingly important in e-commerce scenarios. Current mainstream approaches typically use a two-stage training strategy: first train a product embedding model, and then learn a codebook that maps embeddings to product IDs. This cascaded approach suffers from two major issues: (1) error accumulation-if the embedding model in the first stage produces biased representations, the codebook in the second stage cannot correct these errors, degrading final retrieval performance; and (2) codebook learning relies so
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers →
- PossiblePossibly related (embedding) · 46%VectorStore Studio →
- PossiblePossibly related (embedding) · 46%We’ve got a workshop on production retrieval-augmented generation with open models, benchmarked end to end, thought it’d be relevant here [D] →
- PossiblePossibly related (embedding) · 46%Embedex →
- FuzzySimilar title/name (fuzzy) · 84%GoogleCloudPlatform/generative-ai →
“Fuzzy title match (0.92): “Generative Retrieval for E-commerce: Jointly Learning Embedd” ≈ “GoogleCloudPlatform/generative-ai””
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%steven2358/awesome-generative-ai →
“Fuzzy title match (0.73): “Generative Retrieval for E-commerce: Jointly Learning Embedd” ≈ “steven2358/awesome-generative-ai””
