default search action
Fazl Barez
Person information
Refine list
refinements active!
zoomed in on ?? of ?? records
view refined list in
export refined list as
2020 – today
- 2024
- [c8]Michelle Lo, Fazl Barez, Shay B. Cohen:
Large Language Models Relearn Removed Concepts. ACL (Findings) 2024: 8306-8323 - [c7]Michael Lan, Philip Torr, Fazl Barez:
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models. EMNLP 2024: 12576-12601 - [c6]Clement Neo, Shay B. Cohen, Fazl Barez:
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions. EMNLP 2024: 16681-16697 - [c5]Philip Quirke, Fazl Barez:
Understanding Addition in Transformers. ICLR 2024 - [c4]Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schröder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob N. Foerster:
Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AI. ICML 2024 - [c3]Pengyi Li, Jianye Hao, Hongyao Tang, Yan Zheng, Fazl Barez:
Value-Evolutionary-Based Reinforcement Learning. ICML 2024 - [i27]Michelle Lo, Shay B. Cohen, Fazl Barez:
Large Language Models Relearn Removed Concepts. CoRR abs/2401.01814 (2024) - [i26]Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam S. Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul F. Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, Ethan Perez:
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. CoRR abs/2401.05566 (2024) - [i25]Philip Quirke, Clement Neo, Fazl Barez:
Increasing Trust in Language Models through the Reuse of Verified Circuits. CoRR abs/2402.02619 (2024) - [i24]Clement Neo, Shay B. Cohen, Fazl Barez:
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions. CoRR abs/2402.15055 (2024) - [i23]Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schröder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob N. Foerster:
Near to Mid-term Risks and Opportunities of Open Source Generative AI. CoRR abs/2404.17047 (2024) - [i22]Nevan Wichers, Victor Tao, Riccardo Volpato, Fazl Barez:
Visualizing Neural Network Imagination. CoRR abs/2405.06409 (2024) - [i21]Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schröder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Aaron Purewal, Botos Csaba, Fabro Steibel, Fazel Keshtkar, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan Arturo Nolazco, Lori Landay, Matthew Thomas Jackson, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob N. Foerster:
Risks and Opportunities of Open-Source Generative AI. CoRR abs/2405.08597 (2024) - [i20]Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, Evan Hubinger:
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. CoRR abs/2406.10162 (2024) - [i19]Michael Lan, Philip Torr, Austin Meek, Ashkan Khakzar, David Krueger, Fazl Barez:
Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Models. CoRR abs/2410.06981 (2024) - [i18]Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, Fazl Barez:
Towards Interpreting Visual Information Processing in Vision-Language Models. CoRR abs/2410.07149 (2024) - [i17]Tingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen, David Krueger, Fazl Barez:
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning. CoRR abs/2410.08811 (2024) - [i16]Luke Marks, Alasdair Paren, David Krueger, Fazl Barez:
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders. CoRR abs/2411.01220 (2024) - 2023
- [c2]Antonio Valerio Miceli Barone, Fazl Barez, Shay B. Cohen, Ioannis Konstas:
The Larger they are, the Harder they Fail: Language Models do not Recognize Identifier Swaps in Python. ACL (Findings) 2023: 272-292 - [c1]Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, Fazl Barez:
Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark. ACL (Findings) 2023: 11548-11559 - [i15]Fazl Barez, Paul Bilokon, Ruijie Xiong:
Benchmarking Specialized Databases for High-frequency Data. CoRR abs/2301.12561 (2023) - [i14]Fazl Barez, Paul Bilokon, Arthur Gervais, Nikita Lisitsyn:
Exploring the Advantages of Transformers for High-Frequency Trading. CoRR abs/2302.13850 (2023) - [i13]Ondrej Bohdal, Timothy M. Hospedales, Philip H. S. Torr, Fazl Barez:
Fairness in AI and Its Long-Term Implications on Society. CoRR abs/2304.09826 (2023) - [i12]Fazl Barez, Hosien Hasanbieg, Alesandro Abbate:
System III: Learning with Domain Knowledge for Safety Constraints. CoRR abs/2304.11593 (2023) - [i11]Alex Foote, Neel Nanda, Esben Kran, Ioannis Konstas, Fazl Barez:
N2G: A Scalable Approach for Quantifying Interpretable Neuron Representations in Large Language Models. CoRR abs/2304.12918 (2023) - [i10]Antonio Valerio Miceli Barone, Fazl Barez, Ioannis Konstas, Shay B. Cohen:
The Larger They Are, the Harder They Fail: Language Models do not Recognize Identifier Swaps in Python. CoRR abs/2305.15507 (2023) - [i9]Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, Fazl Barez:
Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark. CoRR abs/2305.17553 (2023) - [i8]Alex Foote, Neel Nanda, Esben Kran, Ioannis Konstas, Shay B. Cohen, Fazl Barez:
Neuron to Graph: Interpreting Language Model Neurons at Scale. CoRR abs/2305.19911 (2023) - [i7]Albert Garde, Esben Kran, Fazl Barez:
DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models. CoRR abs/2310.01870 (2023) - [i6]Kayla Matteucci, Shahar Avin, Fazl Barez, Seán Ó hÉigeartaigh:
AI Systems of Concern. CoRR abs/2310.05876 (2023) - [i5]Luke Marks, Amir Abdullah, Luna Mendez, Rauno Arike, Philip H. S. Torr, Fazl Barez:
Interpreting Reward Models in RLHF-Tuned Language Models Using Sparse Autoencoders. CoRR abs/2310.08164 (2023) - [i4]Philip Quirke, Fazl Barez:
Understanding Addition in Transformers. CoRR abs/2310.13121 (2023) - [i3]Michael Lan, Fazl Barez:
Locating Cross-Task Sequence Continuation Circuits in Transformers. CoRR abs/2311.04131 (2023) - [i2]Fazl Barez, Philip H. S. Torr:
Measuring Value Alignment. CoRR abs/2312.15241 (2023) - 2021
- [i1]Cong Wang, Tianpei Yang, Jianye Hao, Yan Zheng, Hongyao Tang, Fazl Barez, Jinyi Liu, Jiajie Peng, Haiyin Piao, Zhixiao Sun:
ED2: An Environment Dynamics Decomposition Framework for World Model Construction. CoRR abs/2112.02817 (2021)
Coauthor Index
manage site settings
To protect your privacy, all features that rely on external API calls from your browser are turned off by default. You need to opt-in for them to become active. All settings here will be stored as cookies with your web browser. For more information see our F.A.Q.
Unpaywalled article links
Add open access links from to the list of external document links (if available).
Privacy notice: By enabling the option above, your browser will contact the API of unpaywall.org to load hyperlinks to open access articles. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Unpaywall privacy policy.
Archived links via Wayback Machine
For web page which are no longer available, try to retrieve content from the of the Internet Archive (if available).
Privacy notice: By enabling the option above, your browser will contact the API of archive.org to check for archived content of web pages that are no longer available. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Internet Archive privacy policy.
Reference lists
Add a list of references from , , and to record detail pages.
load references from crossref.org and opencitations.net
Privacy notice: By enabling the option above, your browser will contact the APIs of crossref.org, opencitations.net, and semanticscholar.org to load article reference information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Crossref privacy policy and the OpenCitations privacy policy, as well as the AI2 Privacy Policy covering Semantic Scholar.
Citation data
Add a list of citing articles from and to record detail pages.
load citations from opencitations.net
Privacy notice: By enabling the option above, your browser will contact the API of opencitations.net and semanticscholar.org to load citation information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the OpenCitations privacy policy as well as the AI2 Privacy Policy covering Semantic Scholar.
OpenAlex data
Load additional information about publications from .
Privacy notice: By enabling the option above, your browser will contact the API of openalex.org to load additional information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the information given by OpenAlex.
last updated on 2024-12-12 21:02 CET by the dblp team
all metadata released as open data under CC0 1.0 license
see also: Terms of Use | Privacy Policy | Imprint