Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
deaux
35 days ago
|
parent
|
context
|
favorite
| on:
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cy...
"Intelligence" being what, math? Coding? Unfortunately there's a billion use cases for LLMs whose performance is not at all captured by the popular benchmarks they're all trying to maxx.
whimsicalism
35 days ago
[–]
if you are relying on a model for a business process, it should be simple enough to benchmark on that process
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: