# Dipankar Sarkar — Applied AI Research and Technical Evaluation > Resolve the claim before committing to the system. Independent researcher offering commissioned > replication studies, benchmark design, model and retrieval comparisons, generated-code correctness > studies and research collaboration. ACM and IEEE member; 149 citations on Google Scholar as of 2026-10-07. ## Identity & Disambiguation Full name: Dipankar Sarkar. IIT Delhi (B.Tech Computer Science, 2003-2007) and Arizona State University (M.S. Computer Science, Cybersecurity specialization, 2020-2022) alumnus. Author of the Fed-Focal Loss paper, the Nginx Web Server Implementation Cookbook (Packt Publishing, 2011), and 25 provisional patent applications filed with the Indian Patent Office (applications, not granted patents). Not to be confused with the Indian film critic or cricketer of the same name. ## Commission or collaborate Study fees are fixed at scoping and do not depend on the result; conflicts are disclosed at intake. ### Overviews - [Agent Evaluation and Coordination Research](https://www.dipankar.cc/research/agent-evaluation/) — Escalation, refusal to fabricate, multi-agent coordination, judge robustness and ranking stability, with papers, artefacts and limits. - [Independent AI Technical Validation and Reproducible Evaluation](https://www.dipankar.cc/research-validation/) — Choose between replication, benchmark design, model comparison, kernel correctness and feasibility studies. - [Generated-Code and GPU-Kernel Correctness Research](https://www.dipankar.cc/research/generated-code/) — Correctness oracles, test-input generation, calibrated tolerances and a 26-op failure corpus for LLM-generated kernels. - [Sponsored Research and Research Collaboration](https://www.dipankar.cc/research-collaboration/) — Sponsored studies, collaborations, consortium work, seminars and open experiments compared by funding, ownership and publication. - [Federated Learning and Privacy Research](https://www.dipankar.cc/research/privacy-and-federated-learning/) — Federated learning under imbalanced data, and the evidence a distributed-ML privacy claim needs: baselines, metrics, threat models. - [Decentralised Systems and Protocol Research](https://www.dipankar.cc/research/decentralised-systems/) — DePIN incentives, composability and MEV: protocol research with stated assumptions, models, experiments and limits. ### Commissioned studies - [AI Benchmark Replication Studies](https://www.dipankar.cc/validation/replication-studies/) — Re-run a published benchmark or paper result under your workload, with raw results and a deviation report. - [AI Benchmark and Evaluation Study Design](https://www.dipankar.cc/validation/benchmark-design/) — Build a pre-specified test for a decision existing benchmarks do not measure, with baselines and failure conditions. - [Independent Model and Retrieval Comparison Studies](https://www.dipankar.cc/validation/model-and-retrieval-comparison/) — Run shortlisted models or retrieval methods on one workload; report uncertainty, input-regime effects and costs. - [Generated GPU Kernel Correctness Studies](https://www.dipankar.cc/validation/generated-kernel-correctness/) — Test generated GPU kernels with operator-aware oracles, calibrated tolerances and adversarial inputs. - [Empirical AI Feasibility Studies](https://www.dipankar.cc/validation/technical-feasibility-studies/) — A bounded experiment with stop criteria fixed in advance, ending in a proceed-or-stop recommendation. ### Sponsored research - [Write a Sponsored AI Research Brief](https://www.dipankar.cc/sponsored-research/study-brief/) — Structure a sponsored AI study around a claim, a decision, a baseline, realistic access and stop criteria. - [Publication, IP and Research Independence](https://www.dipankar.cc/sponsored-research/publication-ip-and-independence/) — A discussion checklist for IP, data handling, publication review, unfavourable results, authorship and conflict disclosure. - [Sponsor an Open Research Experiment](https://www.dipankar.cc/research-sponsorship/) — Fund a public, milestone-based experiment with open methods and artefacts, published whatever the result. ### Collaboration - [Industry-Academic AI Research Collaboration](https://www.dipankar.cc/collaboration/industry-academic-research/) — Map a shared research question, split contributions and settle data and publication dependencies before the work starts. - [Technical Contributions to Research Consortia](https://www.dipankar.cc/collaboration/consortium-technical-contributions/) — Scope a bounded technical work package for a research consortium, with dependencies, deliverables and an honest capacity statement. - [Academic Seminars and Methods Exchanges](https://www.dipankar.cc/academic-seminars/) — Research seminars and methods discussions for labs and seminar series, with an abstract, prerequisite reading and discussion questions. ### Research topics - [Multi-Agent Coordination and Conflicting Work](https://www.dipankar.cc/research/multi-agent-coordination/) — Duplicate effort, conflicting edits and lost writes between coding agents, measured from a coordination log rather than from pull requests. - [RAG Knowledge Conflicts and Input Regimes](https://www.dipankar.cc/research/rag-knowledge-conflicts/) — How conflict detectors for RAG behave across input regimes: five findings from an audit, an input-regime ladder and open questions. - [Efficient Inference and Distillation Evaluation](https://www.dipankar.cc/research/efficient-inference-and-distillation/) — Open research direction: matched protocols, resource accounting and task-specific degradation for compressed and distilled models. - [Privacy Claims and Threat Models in Distributed ML](https://www.dipankar.cc/research/privacy-threat-models/) — The adversary, leakage and trust assumptions a distributed-ML privacy claim must state, and how to design a test for it. - [DePIN Mechanism and Incentive Evaluation](https://www.dipankar.cc/research/depin-mechanism-evaluation/) — Test a DePIN incentive mechanism with stated assumptions, an agent model, sensitivity analysis and a reproducible simulation. - [Composability, MEV and Protocol Mechanisms](https://www.dipankar.cc/research/composability-and-mev/) — Formalise the assumptions behind fairness, MEV and cross-rollup atomicity claims, then search for counterexamples. ### Methods - [Benchmark Leakage and Contamination Checks](https://www.dipankar.cc/methods/benchmark-leakage-and-contamination/) — Overlap audit, written contamination hypotheses and a sensitivity analysis that tests whether the conclusion survives. - [Uncertainty, Repeated Trials and Evaluation Claims](https://www.dipankar.cc/methods/uncertainty-and-repeated-trials/) — Repeated-trial design, cluster bootstrap over items, rank stability and sensitivity comparisons for small performance gaps. - [Agent Failure Taxonomies and Incident Coding](https://www.dipankar.cc/methods/agent-failure-taxonomies/) — Versioned codebook, ambiguity rules, double coding with agreement statistics and annotated examples for agent traces. - [Evaluation Dataset Design and Provenance](https://www.dipankar.cc/methods/evaluation-dataset-design/) — Dataset specification, sampling design with controls, per-item provenance ledger, exclusion log and access restrictions. - [Pre-Specified Evaluation Plans](https://www.dipankar.cc/methods/pre-specified-evaluation-plans/) — Versioned research question, primary outcomes, exclusion and analysis rules, and an amendment log, fixed before results. - [Research Reproducibility Packages](https://www.dipankar.cc/methods/reproducibility-packages/) — Environment manifest, inputs with provenance, per-table scripts, raw per-run outputs and a known-limits statement, with worked examples. ### Free tools - [Evaluation Brief Builder](https://www.dipankar.cc/resources/evaluation-brief-builder/) — Turn an uncertain technical claim into a structured experiment brief in your browser, before any conversation. - [Reproducibility Readiness Checklist](https://www.dipankar.cc/resources/reproducibility-checklist/) — Check which missing artefacts would stop another team rerunning your study, and decide what to fix before release. ### Evidence - [Datasets, Benchmarks and Evaluation Artefacts](https://www.dipankar.cc/datasets-and-benchmarks/) — Public corpora, code and data behind the papers, with licences, version status, access conditions and paper links. ## Technical guides - [Governed AI Agents: Authority, Approval, and Evidence](https://www.dipankar.cc/guides/governed-ai-agents/) — A practical architecture for governing tool-using AI agents with identity, purpose, action tiers, approvals, replayable evidence, and human stop conditions. - [Production AI Agent Reliability: State, Side Effects, and Recovery](https://www.dipankar.cc/guides/production-agent-reliability/) — How to design reliable AI agents using explicit state, durable checkpoints, idempotent effects, replay, evaluation, observability, and operator-owned recovery. - [AI-Generated Code Correctness: Test the Semantics, Not the Vibe](https://www.dipankar.cc/guides/ai-generated-code-correctness/) — Testing LLM-generated CUDA and Triton kernels with operator-aware oracles, adversarial inputs, reproducible failures, and false-positive controls. - [Multi-Agent Software Engineering Before the Pull Request](https://www.dipankar.cc/guides/multi-agent-software-engineering/) — How task claims, append-only coordination events, replica convergence, and pre-PR observability expose duplicate and conflicting work among coding agents. - [Local and Embedded AI Inference Across Language Boundaries](https://www.dipankar.cc/guides/local-inference-systems/) — A practical decision framework for local versus hosted inference, in-process model execution, formats, bindings, hardware backends, resource limits, and offline operation. - [Physical AI and Robotics: From Language Plan to Safe Edge Execution](https://www.dipankar.cc/guides/physical-ai-and-robotics/) — The system boundary between natural-language robot planning, deterministic validation, capability constraints, edge execution, telemetry, recovery, and warehouse learning benchmarks. - [Agent-Compatible Tools: Interfaces for Humans and AI Systems](https://www.dipankar.cc/guides/agent-compatible-tools/) — How to design CLIs, APIs, and MCP tools with discoverable capabilities, typed inputs, stable errors, dry runs, scoped credentials, approvals, and audit records. - [Blockchain and Agent Security: Control the Key, State, and Effect](https://www.dipankar.cc/guides/blockchain-agent-security/) — Design patterns for validator signing policy, atomic cross-rollup actions, compiler boundaries, and paper-first agent-operated market systems. ## Other Sites by Dipankar Sarkar Same person, different engagements: - [dipankar.co](https://www.dipankar.co) — AI enablement, contract AI engineering and AI delivery; when you need something built, fixed, operated or adopted - [dipankar.name](https://www.dipankar.name) — Hands-on technology leadership mandates; when you need someone to own a technology function with organisational authority - [dipankar.org](https://www.dipankar.org) — Board advisory, executive workshops, institutional programmes and speaking; when the output is a leadership decision, a programme or a talk - [desinerd.com](https://www.desinerd.com) — Personal blog since 2007; for informal writing ## Publications - Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance (2026-09-28) — arXiv preprint arXiv:2609.37501. https://www.dipankar.cc/publication/bounded-autonomy-regulated-agentic-ai.md/ - When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions (2026-09-27) — arXiv preprint arXiv:2609.33312. https://www.dipankar.cc/publication/on-device-ml-auction-misalignment.md/ - How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure (2026-09-24) — arXiv preprint arXiv:2609.30074. https://www.dipankar.cc/publication/reproducibility-of-evaluation-conclusions.md/ - A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation (2026-08-14) — KDD 2026 Workshop on Secure and Trustworthy Large Language Models (SeT-LLM), poster. https://www.dipankar.cc/publication/four-axis-llm-as-judge-benchmark.md/ - From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving (2026-08-14) — IJCAI-ECAI 2026 Workshop on Logic and Symbolic Reasoning (LogiSymb), poster. https://www.dipankar.cc/publication/minimal-core-guided-repair.md/ - Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels (2026-06-23) — arXiv preprint arXiv:2607.16228. https://www.dipankar.cc/publication/mixed-precision-tolerance-calibration.md/ - Static PTX Metrics Track Structural Kernel Regressions but Miss Semantic Ones (2026-06-23) — arXiv preprint arXiv:2607.02541. https://www.dipankar.cc/publication/static-ptx-metrics-kernel-regressions.md/ - Test-Input Generation for Tensor Programs: What Actually Finds Kernel Bugs (2026-06-20) — arXiv preprint arXiv:2606.27396. https://www.dipankar.cc/publication/test-input-generation-tensor-programs.md/ - The Correctness Illusion in LLM-Generated GPU Kernels (2026-06-15) — arXiv preprint arXiv:2606.20128. https://www.dipankar.cc/publication/correctness-illusion-llm-gpu-kernels.md/ - Before the Pull Request: Mining Multi-Agent Coordination (2026-06-10) — arXiv preprint arXiv:2606.19616. https://www.dipankar.cc/publication/before-the-pull-request-multi-agent-coordination.md/ - An Input-Regime Audit of Conflict Detection for Retrieval-Augmented Generation (2026-05-15) — Manuscript submitted to the 2nd Workshop on Vector Databases (VecDB @ VLDB 2026). https://www.dipankar.cc/publication/input-regime-audit-rag-conflict-detection.md/ - Epistral Network: Revolutionizing Media Curation and Consumption through Decentralization (2024-02-01) — arXiv preprint arXiv:2402.04881. https://www.dipankar.cc/publication/epistral-network.md/ - Navigating the Knowledge Sea: Planet-scale answer retrieval using LLMs (2024-02-01) — arXiv preprint arXiv:2402.05318. https://www.dipankar.cc/publication/navigating-knowledge-sea.md/ - FairFlow Protocol: Equitable Maximal Extractable Value (MEV) mitigation in Ethereum (2023-12-01) — arXiv preprint arXiv:2312.12654. https://www.dipankar.cc/publication/fairflow-protocol.md/ - Viz: A QLoRA-based Copyright Marketplace for Legally Compliant Generative AI (2023-12-01) — arXiv preprint arXiv:2401.00503. https://www.dipankar.cc/publication/viz-qlora-copyright-marketplace.md/ - Centralized Intermediation in a Decentralized Web3 Economy: Value Accrual and Extraction (2023-11-01) — arXiv preprint arXiv:2311.08234. https://www.dipankar.cc/publication/centralized-intermediation-web3.md/ - Decentralized Deepfake Detection Blockchain Network using Dynamic Algorithm management (2023-11-01) — arXiv preprint arXiv:2311.18545. https://www.dipankar.cc/publication/decentralized-deepfake-detection.md/ - Generalised DePIN Protocol: A Framework for Decentralized Physical Infrastructure Networks (2023-11-01) — arXiv preprint arXiv:2311.00551. https://www.dipankar.cc/publication/generalised-depin-protocol.md/ - Towards Universal Atomic Composability: A Formal Model for Multi-Rollup Environments on Ethereum (2023-11-01) — arXiv preprint arXiv:2311.00422. https://www.dipankar.cc/publication/universal-atomic-composability.md/ - Curriculum generation using Autoencoder based continuous optimization (2021-06-01) — arXiv preprint arXiv:2106.08569. https://www.dipankar.cc/publication/curriculum-generation-autoencoder.md/ - One Shot Audio to Animated Video Generation (2021-02-01) — arXiv preprint arXiv:2102.09737. https://www.dipankar.cc/publication/one-shot-audio-animation.md/ - Catfedavg: Optimising communication-efficiency and classification accuracy in federated learning (2020-11-01) — arXiv preprint arXiv:2011.07229. https://www.dipankar.cc/publication/catfedavg.md/ - Fed-Focal Loss for imbalanced data classification in Federated Learning (2020-11-01) — Workshop on Federated Learning for Data Privacy and Confidentiality in Conjunction with IJCAI 2020. https://www.dipankar.cc/publication/fed-focal-loss.md/ - Nginx 1 Web Server Implementation Cookbook — Packt Publishing (2011) (2011-01-01) — Packt Publishing. https://www.dipankar.cc/publication/nginx-cookbook.md/ - Muzzled: A graphical theme builder for Mozilla (2005-01-01) — Linux Gazette. https://www.dipankar.cc/publication/muzzled-mozilla-theme-builder.md/ ## Patents (Provisional Applications, Indian Patent Office) - A Method and a System for Dynamically Modifying a Virtual World — Provisional Application (Indian Patent Office, 2020-08-07) - A System to Generate and Dynamically Update a Virtual World Based on User Interest — Provisional Application (Indian Patent Office, 2020-08-07) - A Method and System to Recommend a Navigation Path in a Virtual World — Provisional Application (Indian Patent Office, 2020-08-07) - Method and System to Generate Animated Audio-Visual Content — Provisional Application (Indian Patent Office, 2020-07-10) - System and Method for Partitioning a Neural Network Model for Offloading Computational Load — Provisional Application (Indian Patent Office, 2020-07-10) - Method for Converting Static Graphic into Animated Graphic — Provisional Application (Indian Patent Office, 2020-07-10) - Method for Dynamic Content Generation and Device Thereof — Provisional Application (Indian Patent Office, 2020-06-30) - A Method and System for Partitioning a Social Network Group — Provisional Application (Indian Patent Office, 2020-05-06) - A Method and System for Improvising a Pre-Generated Avatar of a User — Provisional Application (Indian Patent Office, 2020-04-24) - A System to Generate an In-Context Link to a Shared Experience in a Chat Session — Provisional Application (Indian Patent Office, 2020-04-22) - System to Generate and Recommend Offline Compatible Interactions to Users — Provisional Application (Indian Patent Office, 2020-04-21) - Methods and Systems for Impersonation and Exploration of a Virtual World — Provisional Application (Indian Patent Office, 2020-04-21) - A System to Alter User Profile and Experience on a Virtual Communication Platform — Provisional Application (Indian Patent Office, 2020-04-21) - Method of Intimating Experience/Activity to a User — Provisional Application (Indian Patent Office, 2020-04-17) - Method and System for Transmitting Messages in Low Bandwidth Environment — Provisional Application (Indian Patent Office, 2020-04-14) - A System and Method for Generating a Video/Animation for Matched Users — Provisional Application (Indian Patent Office, 2020-04-14) - Methods and Systems to Transform a User Recorded Video into a Video Sticker — Provisional Application (Indian Patent Office, 2020-02-24) - System and Method for Generating Animated Visual Appearance of User, Based on Audio Message — Provisional Application (Indian Patent Office, 2020-02-18) - A System and Method for Generating Unified Image on a Messaging Platform — Provisional Application (Indian Patent Office, 2020-02-18) - Systems and Methods for Converting Text to Speech Mimicking User's Voice Tone — Provisional Application (Indian Patent Office, 2020-02-03) - A Method and System for Generating Multiple Expressive Emojis — Provisional Application (Indian Patent Office, 2020-01-06) - A System for Providing Avatars from an Encrypted Image of a User — Provisional Application (Indian Patent Office, 2020-01-06) - A Method and System for Generating Hairstyle Vector — Provisional Application (Indian Patent Office, 2020-01-06) - Natural Language Query Based Search System — Provisional Application (Indian Patent Office, 2020-01-06) - A Method and a System for Determining a Relationship Type Between Users — Provisional Application (Indian Patent Office, 2020-01-06) ## Posts - [Fair Ordering: Taming MEV on Ethereum](https://www.dipankar.cc/post/fair-ordering-mev.md/) — Maximal Extractable Value comes from a validator freedom to order transactions within a block, and every mitigation approach constrains that freedom differently - [Why GPU Kernel Benchmarks Miss Real Bugs](https://www.dipankar.cc/post/gpu-kernel-benchmark-blindspots.md/) — Passing a correctness benchmark and being correct in production are different claims for LLM-generated GPU kernels — the gap comes down to which inputs get samp - [Diagnosing Conflicts in RAG Pipelines](https://www.dipankar.cc/post/rag-conflict-diagnosis.md/) — Why retrieval-augmented generation fails when context disagrees with parametric knowledge, and a checklist for diagnosing which failure mode you are looking at. - [From AutoPrompt to TextGrad: A Prompt Survey](https://www.dipankar.cc/post/automated-prompt-optimization-survey.md/) — A chronological survey of automated prompt optimization 2020–2025: AutoPrompt, APE, OPRO, EvoPrompt, DSPy, TextGrad, PromptAgent, and how to choose between them - [LLM Prompt Compression: A Practitioner Guide](https://www.dipankar.cc/post/llm-prompt-compression-guide.md/) — A practitioner's guide to LLM prompt compression: LLMLingua, GIST Tokens, 500xCompressor, KV-cache methods, and the rate-distortion limits of compressing contex - [Prompt Structuring Techniques for LLMs](https://www.dipankar.cc/post/llm-prompt-structuring-techniques.md/) — A chronological survey of LLM prompt structuring: chain-of-thought, the instruction hierarchy, system prompt design, and evaluation frameworks. - [Comparing LLM Safety Techniques](https://www.dipankar.cc/post/llm-safety-techniques-survey.md/) — A practitioner's survey of LLM safety techniques across OpenAI Harmony, Anthropic Constitutional AI, Google SAIF, Meta Llama Guard, and open-source RLHF framewo - [Tackling Data Imbalance in Federated Learning](https://www.dipankar.cc/post/federated-learning-imbalanced-data.md/) — How Fed-Focal Loss addresses one of the most challenging problems in distributed machine learning: handling imbalanced data across federated clients. - [The AI Copyright Challenge for Generative AI](https://www.dipankar.cc/post/ai-copyright-challenge.md/) — As generative AI transforms content creation, we need new frameworks that respect copyright while enabling innovation. Here's how we can build them. - [The MEV Problem: Fairer Value Distribution](https://www.dipankar.cc/post/mev-ethereum-fairness.md/) — Exploring Maximal Extractable Value (MEV) in Ethereum and why we need better mechanisms for fair value distribution across the ecosystem. - [DePIN: The Future of Physical Infrastructure](https://www.dipankar.cc/post/depin-future-infrastructure.md/) — Why Decentralized Physical Infrastructure Networks (DePIN) represent a fundamental shift in how we build and own critical infrastructure. - [Why Deepfake Detection Needs Decentralization](https://www.dipankar.cc/post/deepfake-detection-decentralized.md/) — As deepfake technology becomes more sophisticated, centralized detection approaches are failing. Here's why we need decentralized solutions. ## Contact - Website: https://www.dipankar.cc - Email: contact@dipankar.cc - ORCID: https://orcid.org/0000-0001-5431-6367 - Google Scholar: https://scholar.google.com/citations?user=t_ikr2UAAAAJ&hl=en - GitHub: https://github.com/sarkar-dipankar (research) / https://github.com/dipankar (Rust CLIs) - LinkedIn: https://www.linkedin.com/in/dipankarsarkar