Automatically translated.View original post

Use which model? This post has the answer.

Bundle of popular LLM models. Where are you?

What model do you have for each job?

.

We bring the model card or system card of each model to NotebookLM. Please give us the summary. We will give you the coordinates under the comment.

And here are text models, mainly coding, but in real life they can also be used.

-

🏠, what is each AI model?

- OpenAI: Home of ChatGPT, it will have any GPT model, such as GPT-5.2, GPT-5 mini, GPT-5.3-Codex, GPT-5.3 Codex Low, GPT-5.3 Codex High, GPT-5.3 Codex Extra High, GPT-5.3 Codex Fast, GPT-5.3 Codex Low Fast, GPT-5.3 Codex High Fast, GPT-5.3 Codex Extra High Fast, which also has a real sub-codex model.

- Anthropic: Claude's home, whether it's Claude Haiku, Claude Opus, Claude Sonnet.

- Google: This house will also be Gemini Flash with Gemini Pro.

- xAI: This one is from Elon Mask's X house, it will be Grok Code Fast 1, Grok 4.20.

Basically, we used the first three, and of course, if you can't imagine, choose auto.

This is in the IDE.

- Github Copilot: Well, it's Microsoft. There's a Raptor mini that Microsoft took the GPT-5 mini for fine-tuned, and Goldeneye took the GPT-5.1-Codex for fine-tuned. Both are used in vs code with Github Copilot specifically, and one place.

- Cursor: He himself has his own unique model as well. It's a Composer name. Composer 2.

-

🤖 what kind of job is each model suitable for Dave?

We threw each model card or system card of the model used in the github copilot to NotebookLM to analyze it, because the document of Github Copilot did not update the new model, so we took four topics there to analyze it.

.

⭐ Fast help with simple or repetitive tasks

Focus on simple tasks, do them quickly, focus on quantities such as text classification, introductory language translation, short retrieval, or general question-and-answer chat. There are small models like Gemini 3.1 Flash-Lite / 3 Flash, GPT-5 nano / GPT-5 mini, Claude Haiku 4.5.

.

⭐ General-purpose coding and writing

Focus on general use. Introduce Gemini 3 Pro, GPT-5.2, Claude Sonnet 4.6, a flagship model suitable for everyday use. Have a very good understanding of context. Can write people's language and computer language fluently and with few errors.

.

⭐ Deep reasoning and debugging.

Focus on complex applications and solve bugs. Introduce frontier models with an intensive thought process. Think before you answer, such as Gemini 3.1 Pro (with Deep Think Mode), GPT-5.4 pro / GPT-5.1-Codex-Max, Claude Opus 4.6.

.

⭐ Multi-model (screenshot, diagram)

Any work on images, diagrams, charts, or multi-models from screenshots or image files is analyzed.

Introduces multi-model models like Gemini 3 Pro / Flash, Claude Opus / Sonnet, GPT-5 that he understands the relationship between text and images very well.

.

✨ Gemini's distinctive features: Gemini's system supports uploading large files up to 2 GB. Whether it's video, audio files, images, or thick PDF documents, you can throw them in together to provide cross-modal AI analysis directly on the chat page.

-

🌐 for Web Developer

You can go to the Arena AI web (or the former LMArena) at Leaderboard in the Code category.

The latest update was on April 14, with over 1 million votes! and 334 models.

1. claude-opus-4-6-thinking

2. Claude-opus-4-6

3. Glm-5.1: This is the open source of the z.ai

4. Claude-sonnet-4-6

5. Claude-opus-4-5-2025 1101-thinking-32k

How do they score or vote? They will have a feature battle. We let AI do something like a cart checkout. They have 2 options for us to choose which one they like. We will not know which model each option comes from until we choose which one we like. The code is a one-page html that has stuffed everything.

-

📱 for Android Developer

The Android team has released Android Bench which model is most suitable for our Android app writing.

The latest update, April 7, adds new OpenAI models like GPT-5.4 and GPT-5.3-Codex.

1.GPT-5.4: 72.4% score

2.Gemini 3.1 Pro Preview: 72.4% score - same score as GPT-5.4, but slightly wider CI range value.

3.GPT-5.3-Codex: 67.7% score

4.Claude Opus 4.6: 66.6% score

5.GPT-5.2-Codex: 62.5% score

6.Claude Opus 4.5: 61.9% score

The score value is calculated based on an average of 10 rounds per 100 Test Cases.

Confidence Interval (CI) is the confidence score, the expected performance range of the model, based on the p-value statistically significant level < 0.0.

Since the score may be somewhat volatile, this CI allows us to know where the model's actual score will be.

So how does he think?

The problem comes from over 38,000 actual Pull Requests on Github, but 100 specific Android technical items such as Jetpack Compose, Coroutine with Flows, Room, Hilt are selected by writing a solution called a patch file.

In addition to running through pretty code, you have to run Unit Test or Instrumentation Test to get a score.

It's all about this. The link is under the comment.

# Includes IT matters # Share IT Trick # lemon 8 Howtoo # programmer # Developer

4/17 Edited to

... Read moreถ้าคุณกำลังเสิร์ชหา “GPT-5.3 Codex Extra High Fast” เพราะลังเลว่าจะเลือกโมเดลนี้ใน Cursor (หรือใน IDE ที่ต่อกับโมเดลตระกูล Codex) ดีไหม ขอสรุปแบบประสบการณ์ใช้งานจริง + แนวทางเลือกให้ตรงงานนะ อย่างแรก ชื่อ “Extra High Fast” อ่านแล้วงงนิดนึง แต่โดยพฤติกรรมการใช้งานมันมักจะสื่อ 2 อย่างพร้อมกันคือ (1) โหมดที่เน้นความเร็วในการตอบ และ (2) ระดับความสามารถ/ความแม่นที่ถูกดันให้สูงกว่ารุ่น Fast ปกติ (โดยเฉพาะงานโค้ด) ดังนั้นมันจะเหมาะกับงานที่ “อยากได้คำตอบเร็ว แต่ยังต้องไว้ใจคุณภาพพอสมควร” เช่น ทำฟังก์ชันย่อย ๆ, refactor โค้ดหลายไฟล์แบบไม่ต้องคิดลึกมาก, เขียนเทสพื้นฐาน, หรือช่วยสรุปโค้ด/อธิบาย context ให้ทีม แล้วควรใช้ตอนไหน? 1) งานโค้ดที่ต้องสปีด: เช่น generate boilerplate, เขียน REST/GraphQL client, สร้าง DTO/Model, แปลง API response, ทำ migration เล็ก ๆ หรือแก้ lint/format แบบเป็นชุด ๆ งานพวกนี้ “เร็ว” สำคัญกว่า “คิดลึก” รุ่น Extra High Fast จะช่วยให้ flow ไม่สะดุด 2) งาน debug ระดับกลาง: ถ้าเป็นบัคที่อ่าน stack trace แล้วพอเดาได้ หรือมี test/failing case ชัด ๆ โมเดลสาย Codex ที่เร็วจะช่วยไล่จุดผิด + เสนอ patch ได้ไว แต่ถ้าเป็นบัคเชิงระบบ/ต้อง reasoning หนัก ๆ (เช่น concurrency, race condition, memory leak ซับซ้อน) แนะนำสลับไปตัวที่เน้น deep reasoning ก่อน แล้วค่อยกลับมาใช้ Fast ช่วยทำแพตช์หลายเวอร์ชัน 3) ใช้กับ Cursor ให้คุ้ม: ใน Cursor เวิร์กโฟลว์ที่ผมใช้คือ - “Chat/Ask” ใช้โมเดลเร็ว (Fast) เพื่อถาม-ตอบสั้น ๆ เช่น ขอแนวทาง, ขอ snippet, ขอเช็คลิสต์ - “Agent/Composer” (งานแก้หลายไฟล์) ถ้าเป็นงานเปลี่ยนโครงสร้างใหญ่ ๆ ค่อยขยับไปโมเดลที่นิ่งกว่า แต่ถ้าเป็นงานแก้ตามสเปกชัด ๆ (เช่นเพิ่ม field, เปลี่ยนชื่อ, สร้าง endpoint) Extra High Fast ทำได้ดีและประหยัดเวลา ทริคเล็ก ๆ ให้ได้ผลลัพธ์ดีขึ้น (โดยเฉพาะเวลาใช้ใน IDE): - แนบ context ให้พอดี: แปะโค้ดเฉพาะไฟล์/ฟังก์ชันที่เกี่ยวข้อง + error message + expected behavior อย่าโยนทั้ง repo ถ้าไม่จำเป็น - สั่งให้ “เสนอ patch แบบ minimal”: บอกให้แก้เท่าที่จำเป็นและอธิบายเหตุผล 2-3 บรรทัด จะลดการแก้เกิน - บังคับให้เขียนเทส/คำสั่งรัน: เช่น “เพิ่ม unit test 1 เคส + บอกคำสั่งรัน” จะช่วยกัน regressions สุดท้าย ถ้าคุณยังเลือกไม่ถูกจริง ๆ ให้ยึดหลักง่าย ๆ: งานซ้ำ ๆ และต้องเร็ว → เลือก Fast, งานทั่วไปเขียนโค้ด/เขียนเอกสาร → เลือก flagship, งานคิดหนัก/ดีบักยาก → เลือก deep reasoning แล้วค่อยให้ Fast ช่วยทำโค้ดให้เสร็จเร็ว ๆ