Göktuğ Düşünen PRO
AI & ML interests
Head of AI @ Werea · AI/ML Engineer
AI/ML Engineer building end-to-end intelligent systems across LLMs, NLP, Computer Vision, RAG, Retrieval, Agents, and Applied AI.
From data and model development to fine-tuning, evaluation, optimization, deployment, and production AI systems.
Recent Activity
posted an update 2 days ago
🇹🇷 We trained a 1B OCR model specifically for Turkish enterprise documents.
**Werea-DocOCR-1B v2**
The result surprised us:
LightOnOCR-2 base → **64.2% CER**
Werea-DocOCR v1 → **~8.1% CER**
Werea-DocOCR v2 → **0.15% CER** 🚀
Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.
📄 12 Turkish enterprise document types
🧪 12,960 synthetic training pages
📱 Digital + scanned + phone photos
📊 Tables → structured Markdown
⚙️ Full-parameter fine-tuning
🖥️ Trained on a single RTX 3090
It handles:
• e-Invoices
• rental contracts
• bank receipts
• payroll documents
• insurance policies
• vehicle documents
• official correspondence
• trade registry documents
• SGK-style tables
• and more.
**Model 🤗**
https://huggingface.co/Werea-co/Werea-DocOCR-1B
**Dataset 📚**
https://huggingface.co/datasets/Werea-co/werea-tr-doc-ocr-enterprise-v2
**Werea 🇹🇷**
https://huggingface.co/Werea-co
We're building open AI models from Türkiye.
This is just the beginning.
#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision
reacted to jasoncorkill's post with ❤️ 3 days ago
Most public benchmarks collapse model performance into one broad preference signal.
That makes it hard to understand which capabilities differentiate between models. It's also almost impossible to inspect the evidence behind it. So @RapidataAI is releasing Benchmark.AI.
We started with an SVG generation benchmark including 42 models, 500 prompts, 1.9M+ human judgements, 300K+ match-ups.
We evaluate models separately on Preference, Alignment and Coherence, while making the prompts, outputs, match-ups and methodology public.
Full dataset: https://huggingface.co/datasets/Rapidata/svg-benchmark
Full benchmark: https://www.benchmark.ai/svg
Methodology feedback and benchmark suggestions very welcome!
reacted to theirpost with 🔥 3 days ago
🇹🇷 Introducing T3 Gemstone — Edge AI & Cybersecurity Models from Türkiye
We've been building an open AI ecosystem focused on practical models that can run closer to the edge — not only in large datacenters.
Today, I'm introducing T3 Gemstone, a growing family of compact AI models built around edge inference, cybersecurity and computer vision.
💎 T3 Gemstone currently includes:
* NanoSOC Gemstone 2B — GGUF
* NanoSOC Gemstone 4B — GGUF
* Gemstone Person/Object Detector Nano
* Edge-focused AI experiments and deployments
The goal is simple:
Build smaller, practical and open AI systems that can actually run on constrained hardware.
This is part of a much larger open-source effort we're building from Türkiye across LLMs, cybersecurity, computer vision, retrieval, speech and edge AI.
There is much more coming.
🤗 Explore my models, datasets and demos:
https://huggingface.co/GoktugD
🛡️ NanoSOC:
https://huggingface.co/Werea-co/Werea-NanoSOC-8B
Feedback, benchmarks, collaborations and contributions are very welcome.
If you're interested inopen-source AI, Turkish AI research, edge AI or cybersecurity models, follow the journey.
We're just getting started. 🇹🇷
#AI #OpenSource #HuggingFace #LLM #EdgeAI #Cybersecurity #ComputerVision #TurkishAI #MachineLearning