Dipankar Sarkar's picture
🏗️ Building on HF

Dipankar Sarkar PRO

dipankarsarkar

AI & ML interests

Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.

Recent Activity

liked a dataset about 4 hours ago
RiverRider/srt-cxr14-frozen-probe
repliedto RiverRider's post about 4 hours ago
A 339 KB linear probe on frozen features beats the fine-tuned baseline on ChestX-ray14. Linear(5376, 14) on frozen google/gemma-4-31B-it hidden states. No fine-tuning, no radiology pretraining, no augmentation. All 112,120 images, official test_list.txt. Wang et al. 2017, ResNet-50 fine-tuned end to end 0.7451 this probe, frozen backbone + linear head 0.7590 view-position only (shortcut baseline) 0.5896 shuffled labels (refit floor) 0.5002 Ahead on 12 of 14 findings. The comparison is split-matched, and that took care to get right. The number everyone quotes, CheXNet's 0.8414, is on a different test set: their own random 70/10/20 partition, not the official list. Do not compare 0.7590 to it. The matched row is from Wang's v5 appendix, added specifically to report the published split. I had this wrong in our own code for a day, quoting a cross-split reference as a head-to-head, which is the error worth not repeating in public. Three controls, because a bare AUROC here is not interpretable. Shuffled labels catch leakage. View-only catches the shortcut, since portable AP films are taken of sicker patients, and it is folded, because Hernia's raw view-only of 0.3436 is really 0.6564 of shortcut once flipped. Intervals resample patients and not images, since the test split is 25,596 films from 2,797 patients. Banked negatives are on the card too. Max-pooling and top-16 pooling were predicted to help focal findings and did the opposite, costing 0.0537 and 0.0225. Readout depth barely matters, 0.7600 to 0.7605. Scope: detection, not early detection. Research artifact, not a diagnostic device. The backbone never runs in the demo. What ships is the reading. Space: https://huggingface.co/spaces/RiverRider/srt-cxr14-probe Model: https://huggingface.co/RiverRider/srt-cxr14-linear-probe Data + states: https://huggingface.co/datasets/RiverRider/srt-cxr14-frozen-probe
repliedto kanaria007's post about 4 hours ago
✅ Article highlight: Benchmark Publication Without Governance Inflation (art-60-274, v0.1) TL;DR: This article argues that a benchmark result is not a governance maturity claim. A score may be real, reproducible, and worth publishing—and still say nothing by itself about safety, deployability, assurance, institutional quality, or platform maturity. 274 treats benchmark publication as a discipline of comparability, disclosure, lifecycle limits, and anti-inflation. Read: https://huggingface.co/datasets/kanaria007/agi-structural-intelligence-protocols/blob/main/article/60-supplements/art-60-274-benchmark-publication-without-governance-inflation.md Why it matters: • prevents measured results from being inflated into safety or maturity claims • separates historical results from current comparability • makes scope, freshness, omissions, and unsupported readings visible • allows honest publication without requiring full platform assurance • treats narrower wording as trust discipline, not underselling What’s inside: • the publication triad: comparability, disclosure, and anti-inflation • bounded publication outcomes such as PUBLISHABLE, PUBLISHABLE_WITH_LIMITS, NOT_COMPARABLE, and NOT_PUBLISHABLE • benchmark publication profiles • comparability disclosure notes • public non-claims registers • inflation checklists for result-to-maturity, comparison-to-assurance, historical-to-current, and wording inflation Key idea: Do not say: “this system scored well, therefore it is mature, safe, or ready to deploy.” Say: “this result was observed under this benchmark and comparability frame, remains valid within these lifecycle and disclosure limits, and does not support these broader governance claims.” Better benchmark publication is not a louder score. It is a result that is harder to overread.
View all activity

Organizations

Skelf Research's profile picture Neul Labs's profile picture Cognisoc's profile picture Incredlabs's profile picture