Search: [llm]

[2505.23836] Large Language Models Often Know When They Are Being Evaluated

If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations,...

paper · llm

June 4, 2025 at 9:32:16 PM EDT * · permalink

·

https://www.arxiv.org/abs/2505.23836

SWE-smith

Creating training data for software engineering agents is difficult. Until now.

Introducing SWE-smith: Generate 100s to 1000s of task instances for any GitHub repository.

We've generated 50k+ task instances for 128 popular GitHub repositories, then trained our own LM for SWE-agent.

The result? SWE-agent-LM-32B achieve 40% pass@1 on SWE-bench Verified.

Now, we've open-sourced everything, and we're excited to see what you build with it!

Check out the tutorial below to generate 100 task instances for any GitHub repository in 10 minutes.

llm

May 16, 2025 at 10:01:37 PM EDT * · permalink

·

https://swesmith.com/

Paper page - LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities

The success of Large Language Models (LLMs) has sparked interest in various agentic applications. A key hypothesis is that LLMs, leveraging common sense and Chain-of-Thought (CoT) reasoning, can effectively explore and efficiently solve complex domains. However, LLM agents have been found to suffer from sub-optimal exploration and the knowing-doing gap, the inability to effectively act on knowledge present in the model. In this work, we systematically study why LLMs perform sub-optimally in decision-making scenarios. In particular, we closely examine three prevalent failure modes: greediness, frequency bias, and the knowing-doing gap. We propose mitigation of these shortcomings by fine-tuning via Reinforcement Learning (RL) on self-generated CoT rationales. Our experiments across multi-armed bandits, contextual bandits, and Tic-tac-toe, demonstrate that RL fine-tuning enhances the decision-making abilities of LLMs by increasing exploration and narrowing the knowing-doing gap. Finally, we study both classic exploration mechanisms, such as epsilon-greedy, and LLM-specific approaches, such as self-correction and self-consistency, to enable more effective fine-tuning of LLMs for decision-making.

paper · llm

April 27, 2025 at 4:29:07 PM EDT * · permalink

·

https://huggingface.co/papers/2504.16078

The First LLM – Jonathon Belotti [thundergolfer]

A tracing of the history of GPT-1 and its predecessors.

llm

March 30, 2025 at 3:29:19 PM EDT * · permalink

·

https://thundergolfer.com/blog/the-first-llm

Gitingest

Replace 'hub' with 'ingest' in any GitHub URL for a prompt-friendly text.

github · ai · llm

February 13, 2025 at 8:51:09 PM EST * · permalink

·

https://gitingest.com/

I can now run a GPT-4 class model on my laptop

llm

December 9, 2024 at 2:48:50 PM EST * · permalink

·

https://simonwillison.net/2024/Dec/9/llama-33-70b/

OpenAI's Whisper model is reportedly 'hallucinating' in high-risk situations | Tom's Guide

A new report reveals OpenAI's audio transcription tool, Whisper, has recorded consistent "hallucinations", according to multiple studies.

llm

October 28, 2024 at 4:42:05 PM EDT * · permalink

·

https://www.tomsguide.com/ai/openais-whisper-model-is-reportedly-hallucinating-in-high-risk-situations

Google's rumored Gemini 2.0 launch in December could support LLM stagnation thesis

Google is gearing up to unveil its latest AI language model, Gemini 2.0, in December, according to insider sources from The Verge.

Another indication of the plateau thesis: OpenAI has just confirmed that a new model, internally considered as a potential successor to GPT-4, will not be released this year, despite looming competition from Google Gemini 2.0.

Similarly, Anthropic is rumored to have put a previously announced version 3.5 of its flagship Opus model on hold due to a lack of significant progress, instead focusing on an improved version of Sonnet 3.5 that emphasizes agent-based AI.

google · llm

October 26, 2024 at 2:46:43 PM EDT * · permalink

·

https://the-decoder.com/googles-rumored-gemini-2-0-launch-in-december-could-support-llm-stagnation-thesis/

Introducing the Open FinLLM Leaderboard

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

llm · finance

October 6, 2024 at 1:37:00 PM EDT * · permalink

·

https://huggingface.co/blog/leaderboard-finbench

[2203.14465] STaR: Bootstrapping Reasoning With Reasoning

Generating step-by-step "chain-of-thought" rationales improves language model performance on complex reasoning tasks like mathematics or commonsense question-answering.

ml · llm · paper

September 14, 2024 at 7:20:47 PM EDT * · permalink

·

https://arxiv.org/abs/2203.14465

THE HYBRID FORECAST OF S&P 500 VOLATILITY ENSEMBLED FROM VIX, GARCH AND LSTM MODELS

hybrid LSTM models, significantly outperform the traditional GARCH models

finance · paper · llm

September 9, 2024 at 10:20:59 AM EDT * · permalink

·

https://www.wne.uw.edu.pl/application/files/4417/1949/0286/WNE_WP449.pdf

[2409.01666] In Defense of RAG in the Era of Long-Context Language Models

llm · ai · paper

September 4, 2024 at 10:27:45 PM EDT * · permalink

·

https://arxiv.org/abs/2409.01666

Anthropic launches Claude Enterprise plan to compete with OpenAI | TechCrunch

Anthropic is launching a new subscription plan for its AI chatbot, Claude, catered toward enterprise customers that want more administrative controls and

llm · ai

September 4, 2024 at 4:57:35 PM EDT * · permalink

·

https://techcrunch.com/2024/09/04/anthropic-launches-claude-enterprise-plan-to-compete-with-openai

Faith and Fate: Transformers as fuzzy pattern matchers – Answer.AI

llm · to_read

August 27, 2024 at 3:02:57 PM EDT * · permalink

·

https://www.answer.ai/posts/2024-07-25-transformers-as-matchers.html

Anthropic's new prompt caching will save developers a fortune | VentureBeat

Anthropic's prompt caching lets users save prompts and call these up for later sessions with additional context for a lower price.

llm

August 15, 2024 at 5:14:33 PM EDT * · permalink

·

https://venturebeat.com/ai/anthropics-new-claude-prompt-caching-will-save-developers-a-fortune/

Unveiling Hermes 3: The First Fine-Tuned Llama 3.1 405B Model is on Lambda’s Cloud

We’re excited to offer the AI/ML community free access to Hermes 3 through Lambda’s new Chat Completions API, fully compatible with the OpenAI API. It provides endpoints for creating completions, chat completions and listing models.

llm

August 15, 2024 at 2:34:37 PM EDT * · permalink

·

https://lambdalabs.com/blog/unveiling-hermes-3-the-first-fine-tuned-llama-3.1-405b-model-is-on-lambdas-cloud

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

llm

July 18, 2024 at 5:34:44 PM EDT * · permalink

·

https://livecodebench.github.io/

Slack Combines ASTs with Large Language Models to Automatically Convert 80% of 15,000 Unit Tests - InfoQ

Slack's engineering team recently published how it used a large language model (LLM) to automatically convert 15,000 unit and integration tests from Enzyme to React Testing Library (RTL). By combining

llm

June 12, 2024 at 10:37:54 AM EDT * · permalink

·

https://www.infoq.com/news/2024/06/slack-automatic-test-conversion/