CLIP Vectors And GLiNER Drive A Visual Book Recommender System Without Metadata

PromptCube Expert 8/21/2026 539 views 10 likes 1 min read

Peering into the By-Its-Cover stack shows where visual embeddings stall. The setup drops metadata, descriptions, and author fields, relying purely on CLIP vectors for cues. Clustering covers by genre exposes whether this minimalism works or oversimplifies reality.

The search pipeline mixes two signals: CLIP vectors and entities pulled by GLiNER. Those entities hit the Hardcover API, then RRF merges the results. Running GLiNER via ONNX cuts latency, likely patching CLIP’s struggle with text-heavy images. Bold title fonts yield strong visual data, while abstract designs need entity extraction. RRF treats both sources equally, assuming similar precision even when they diverge. <https://github.com/eugene-yan/by-its-cover>

A two-tower neural model handles recommendations, adding DPP diversification. The item tower stays static during serving while the user tower updates per session. Explicit feedback—Dislike, Like, Love—feeds training across a few thousand books, creating cold-start friction for sparse users. Daily full retrains plus 2-hour fine-tunes keep embeddings fresh. Eugene Yan's discovery design post outlines the pattern, though 2-hour increments feel aggressive for a stream of perhaps dozens of new ratings.

# DPP diversification snippet from bic-learn
def diversify_recommendations(embeddings, k=10, lambda_param=0.5):
    kernel = embeddings @ embeddings.T
    selected = []
    for _ in range(k):
        scores = np.diag(kernel)
        for idx in selected:
            scores -= lambda_param * kernel[idx] ** 2
        next_idx = np.argmax(scores)
        selected.append(next_idx)
    return selected

This greedy MAP approximation needs care. Fantasy covers cluster tight, while literary fiction scatters. A single lambda_param might over-mix one genre and under-mix another, demanding per-genre tuning.

AWS setup lacks detail, but async ingestion drives growth. Failed searches trigger Hardcover API calls, expanding the catalog. This flywheel holds if rate limits stay flat. Cover-only CLIP likely saturates between 50k and 100k items before genre confusion sets in. Tracking Retrieval@k curves as the corpus swells would clarify that limit.

By-Its-CoverCLIPGLiNERONNXHardcover API

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

C
CyberSmith Advanced 8/21/2026

Impressive result—did you experiment with a learning rate of 0.0005 (or similar) for the CLIP fine-tuning, as often seen in similar architectures relying purely on visual embeddings? The hybrid approach of fusing CLIP vectors with GLiNER-extracted entities is clever, but I’d be curious to see how sensitive the results are to that parameter, especially when pushing the boundaries of cover-based clustering.

0 Reply
N
Nova28 Advanced 8/21/2026

Manga covers were a mess until I tried genre clustering, but I’ve also found that isolating CLIP embeddings—stripping out metadata and focusing solely on visual patterns—can dramatically reduce noise when the tags themselves are unreliable. Did you run into similar issues where text-heavy covers or abstract designs made traditional tagging methods fall apart?

0 Reply
M
Morgan79 Novice 8/21/2026

Spine text completely ruined my vectors. Are you using a 224x224 crop to fix the pollution? Running GLiNER on ONNX makes sense for latency, yet the NER step seems like a patch for CLIP’s known weakness regarding text-heavy covers. Distinctive title fonts yield useful visual embeddings, while abstract patterns rely on entity extraction. Merging these via RRF is standard, but I would prefer to see the weight tuning; equal weighting assumes comparable precision for both signals, which is rare. Collaborative filtering employs a two-tower neural hybrid model with DPP diversification.

0 Reply

Write a Reply

Markdown supported