Skip to content
View goktugozkanmd's full-sized avatar

Block or report goktugozkanmd

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
goktugozkanmd/README.md

Göktuğ Özkan, MD

Internal medicine physician building open evaluation systems for medical AI safety and open source AI.

I created MedFailBench, an open clinical safety evaluation project with 44 clinician reviewed cases across 21 clinical domains. I also maintain the Türkiye Clinical AI Evidence Atlas, a bilingual public evidence map with a reproducible method, a 63 field schema, and a validated 25 record pilot.

Repositories

The Repositories tab mixes original projects with contribution forks. Use the lists:

  • Flagship — original public work (MedFailBench, TurKMedBench, evidence atlas, and related projects)
  • Forks — copies used for upstream pull requests, not original projects

Flagship work

Selected merged open source contributions

Active benchmark integrations

Elsewhere

Popular repositories Loading

  1. medical-ai-failure-atlas medical-ai-failure-atlas Public

    MedFailBench exposes critical safety boundary failures in medical AI through clinician built, reproducible open source evaluation.

    Python 3 1

  2. tr-ai-card-radar tr-ai-card-radar Public

    Audit Hugging Face model/dataset cards for Turkish AI resources and write small, reproducible metadata reports. Clinician-led, open-source. No model ranking, no legal/clinical claims.

    Python

  3. lighteval lighteval Public

    Forked from huggingface/lighteval

    Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

    Python

  4. lm-evaluation-harness lm-evaluation-harness Public

    Forked from EleutherAI/lm-evaluation-harness

    A framework for few-shot evaluation of language models.

    Python

  5. inspect_ai inspect_ai Public

    Forked from UKGovernmentBEIS/inspect_ai

    Inspect: A framework for large language model evaluations

    Python

  6. trust-safety-evals trust-safety-evals Public

    Forked from The-AI-Alliance/trust-safety-evals

    The AI Alliance project to define a reference stack for AI model and system evaluation, with evaluations, benchmarks, and leaderboards.

    Makefile